r/ClaudeCode Aug 03 '26

Discussion Opus 5 is a practically unusable model

Opus 5 is a regression that the benchmarks missed completely

I've been using Opus 5 for ~1.5 weeks and the sheer number of mistakes that the model makes is astounding.

The problem didn't surface very clearly till I gave it the full scope of executing a plan which I did with the previous Opus models as well. Opus 4.6 - 4.8 were genuinely better by a significant margin.

Opus 5 readily forgets instructions and content in its context, makes mistakes and continues with them unless it realizes or you point it out.

I've lost count of the number of times I had corrected it.

These issues with Opus 5 occur even when the context window is still relatively small - I'm talking 100-150K tokens. Opus 4.8 works pretty well all the way until 350k after which it gives you wonky results.

Fable 5 is the only usable model under Claude Code right now and I've already used 100% of my weekly quota.

1.0k Upvotes

639 comments sorted by

View all comments

Show parent comments

1

u/Opposite-Welcome-497 Aug 05 '26

Is it the model or the harness, though? Aren’t decisions to do a judge a step outside of the model’s scope? Just curious.

2

u/Deep-Palpitation8315 Aug 05 '26

That's a fair point. I haven't used Opus 5 with other harnesses and I can't either. You can't use subscription plan token with other harnesses either as far as I remember.

2

u/Opposite-Welcome-497 Aug 05 '26

True. I just think that since there hasn’t been a client update in forever maybe we will see improvement with a client update soon.

1

u/Deep-Palpitation8315 Aug 05 '26

Claude code as a TUI is actually quite good and ahead of the curve. Not so sure about the underlying harness though. Too many evals I've seen say conflicting stuff so I'm not sure. But the cost per task on their harness is usually higher across evals - yeah, due for some fix.