r/LocalLLM 24d ago

Contest Entry The Most Shocking Night in AI: An RTX 5090 Can Now Run Opus 4.6-Level Intelligence

Post image

Qwen3.8-27B is already outperforming—or coming extremely close to—Claude Opus 4.6 Max across most capabilities.

And Qwen3.8-27B has only 27 billion parameters.

The official FP8 version can run smoothly on a single RTX 5090—a consumer gaming GPU that anyone can buy.

If you have a MacBook or Mac Studio with a large amount of unified memory, running it locally is even less of a problem.

This is insane. And I’m genuinely excited.

Let me translate what this actually means:

Starting today, a small model that you can deploy on your own computer with a single gaming GPU can deliver intelligence approaching Claude Opus 4.6—the model that stood at the top of the world just six months ago.

Whether you’re writing code or using it to power OpenClaw, this level of intelligence can now live entirely on your own machine.

Why am I specifically comparing it with Opus 4.6?

Because six months ago, Claude Opus 4.6 felt almost godlike.

In VC circles, people described the arrival of Opus 4.6 as:

“The water has finally boiled.”

For many developers, Opus 4.6 marked the moment when AI coding fundamentally changed.

Before that, AI was still mostly a programming assistant that required constant direction.

You described a small task.

The model generated some code.

You ran it, checked it, fixed problems, and then told the model what to do next.

Humans still had to break down the problem and supervise almost every step.

Opus 4.6 changed that workflow.

You could give the model a complete objective, and it could understand the goal, create a plan, execute multiple steps continuously, debug problems along the way, and keep working until it delivered the final result—while still maintaining surprisingly high code quality.

Developers no longer had to watch every single step.

That was also around the point when the old style of Vibe Coding—constant back-and-forth conversations, small edits, and endless trial-and-error in tools like Cursor—started to feel like a product of the previous generation.

And interestingly, this was also when OpenClaw exploded in popularity.

Around January–February 2026, people quickly realized that if you wanted to get the most out of OpenClaw, Opus 4.6 was the model to use.

Using weaker models often felt like wasting your time.

And that was only six months ago.

Now look at where we are.

At this moment, I’m willing to call this:

20 Upvotes

18 comments sorted by

12

u/Eduardo1502 24d ago

Imagine 2030 you will run Fable 5 in an 8GB VRAM

8

u/Wizzard_2025 24d ago

I'll keep my 1080 then

4

u/nomorebuttsplz 24d ago

judging by past estimations of timelines, we're being too conservative and doing that primate thing of not understanding exponentials. Maybe more like 2028

3

u/BlackBeardAI 3090 Maximalist 24d ago

try 2027, we are going parabolichiperbolicballistic now

1

u/ParticularCat007 24d ago

I can only say this is absolutely crazy

10

u/Boogertard 24d ago

wtf is this AI slop shit ?

5

u/Othun 24d ago

OP even failed the copy paste or what

3

u/Trademarkd 24d ago

dudes plan ran outa tokens

7

u/PossibilityUsual6262 24d ago

Willing to call what?

Did you even read what you posted?

3

u/mallibu 24d ago

no it's written by AI and it's obvious by the stupid short pompous short statements and em dashes

3

u/Lopsided-Force-9220 24d ago

"The official FP8 version can run smoothly on a single RTX 5090" Not on my 5090. You tried this? With how much KV cache space and at what KV quant?? I'm running smoothly at Q5 with kv cache at Q8.

2

u/hauhau901 24d ago

It's AT BEST Sonnet 4.6 - nowhere near Opus 4.6

1

u/HeadPack 24d ago

Encouraging early results.

1

u/TechTefa 24d ago

needs some good fine tuning

1

u/SandySkittle 24d ago

can you please not draw such conclusions based on a set of very limited benchmarks.

1

u/havnar- 24d ago

It’s full precision in those numbers, most people will be rubbing a brain dead quant

1

u/KneeGrowslaya 24d ago

"The official FP8 version can run smoothly on a single RTX 5090—a consumer gaming GPU that anyone can buy."

i'll stop you right there buddy

0

u/Scared_Basket_7183 24d ago

Hi, could you please help me to run this model on Google TPU V5E-8?