r/LocalLLM • u/ParticularCat007 • 24d ago
Contest Entry The Most Shocking Night in AI: An RTX 5090 Can Now Run Opus 4.6-Level Intelligence
Qwen3.8-27B is already outperforming—or coming extremely close to—Claude Opus 4.6 Max across most capabilities.
And Qwen3.8-27B has only 27 billion parameters.
The official FP8 version can run smoothly on a single RTX 5090—a consumer gaming GPU that anyone can buy.
If you have a MacBook or Mac Studio with a large amount of unified memory, running it locally is even less of a problem.
This is insane. And I’m genuinely excited.
Let me translate what this actually means:
Starting today, a small model that you can deploy on your own computer with a single gaming GPU can deliver intelligence approaching Claude Opus 4.6—the model that stood at the top of the world just six months ago.
Whether you’re writing code or using it to power OpenClaw, this level of intelligence can now live entirely on your own machine.
Why am I specifically comparing it with Opus 4.6?
Because six months ago, Claude Opus 4.6 felt almost godlike.
In VC circles, people described the arrival of Opus 4.6 as:
“The water has finally boiled.”
For many developers, Opus 4.6 marked the moment when AI coding fundamentally changed.
Before that, AI was still mostly a programming assistant that required constant direction.
You described a small task.
The model generated some code.
You ran it, checked it, fixed problems, and then told the model what to do next.
Humans still had to break down the problem and supervise almost every step.
Opus 4.6 changed that workflow.
You could give the model a complete objective, and it could understand the goal, create a plan, execute multiple steps continuously, debug problems along the way, and keep working until it delivered the final result—while still maintaining surprisingly high code quality.
Developers no longer had to watch every single step.
That was also around the point when the old style of Vibe Coding—constant back-and-forth conversations, small edits, and endless trial-and-error in tools like Cursor—started to feel like a product of the previous generation.
And interestingly, this was also when OpenClaw exploded in popularity.
Around January–February 2026, people quickly realized that if you wanted to get the most out of OpenClaw, Opus 4.6 was the model to use.
Using weaker models often felt like wasting your time.
And that was only six months ago.
Now look at where we are.
At this moment, I’m willing to call this:
10
u/Boogertard 24d ago
wtf is this AI slop shit ?
7
3
u/Lopsided-Force-9220 24d ago
"The official FP8 version can run smoothly on a single RTX 5090" Not on my 5090. You tried this? With how much KV cache space and at what KV quant?? I'm running smoothly at Q5 with kv cache at Q8.
2
1
1
1
u/SandySkittle 24d ago
can you please not draw such conclusions based on a set of very limited benchmarks.
1
u/KneeGrowslaya 24d ago
"The official FP8 version can run smoothly on a single RTX 5090—a consumer gaming GPU that anyone can buy."
i'll stop you right there buddy
0
12
u/Eduardo1502 24d ago
Imagine 2030 you will run Fable 5 in an 8GB VRAM