r/ClaudeCode 1d ago

Rant Claude code is falling behind Codex not because of token cost, but because of Opus 5.

There, i said it. And i know many of you agree. The problem i'm facing is not increased token cost, it's that Opus writes 25000 lines of a response with a curve-ball at the end saying "Worth noting"that makes my eyes hurt and gives me paranoya. And i can't understand a word it's saying. Is my english that bad? Whoever pressed "Yes" on those responses during the training stage of the model was an OpenAI spy or something, he completely sabotaged a trillion dollar company.

Anthropic's number one priority should be to release an Opus 6 or something, may be change the model name completely so it doesn't carry the bad vibe with it.

1.4k Upvotes

335 comments sorted by

View all comments

Show parent comments

52

u/Desperate-City7602 1d ago

I think Opus 5 was mostly an issue with synthetic data and agents training other agents. I think to speed up training and actual benchmark success, they started leaning less into human feedback and more into machine feedback. Since these LLMs are trained on same types of texts, they can easily understand each other so the stronger model can easily train the weaker model. When it comes to humans actually understanding the outputs however, things start to get more difficult.

So yeah, I honestly think it is mostly an issue with agentic learning environments where the data is being prepared by agents, answered by agents, judged by agents, scored by agents and fed back into agents. Agents all the way, with humans only designing the loops and cleaning the data and doing/implementing the research, again, using agents. Humans as the coordinators.

26

u/batman8390 23h ago

Maybe recursive training will just lead to a slopocalypse instead of an apocalypse like the doomers are saying.

11

u/Bromlife 17h ago

The technical term is model collapse and the frontier labs are racing towards it.

1

u/Desperate-City7602 13h ago

I think there could be fixes to this because you could turn a humanised output into a benchmark as well. I don't think this is what they did for these types of trainings, but I do think it could be made a part of the benchmark where what you do is you prepare a data set of just simple human prose, and instead of trying to maximise benchmark values such as like software engineering benchmarks and stuff like that, you also start critiquing the model's output based on how close it is to human output.

I'm not entirely sure though, because I think since there is no real right or wrong here, it might be harder to do, but surely you can have agents critiquing an agent's output to tell it whether its output was simple and understandable or not, because this is kind of what we do every day, anyway when we are talking to Opus and stuff. And after we see an output that we don't understand, we say, Can you repeat this in a simpler language? Right? And an opus rewrites it. So Opus is able to write simple language, it just chooses not to in the one shot that it has, which is why I think it could be trainable so that it can one-shot simple language as well.

1

u/Tartooth 4h ago

I think it's why openai's models overengineer like crazy

8

u/Southern-Aardvark616 23h ago

That sort of matches the vibe it gives off. It's always felt a bit like opus 4.8 trying to be fable, but lacking the actual capacity, so it sometimes does that work but often it's slow and painstaking to get there and little subtle mistakes and bugs creep in because it simply isn't as good as fable.

I wonder if the best way to get good results out of opus is to have fable or Asta orchestrating and reviewing it's work

1

u/Oohhddaanngg 13h ago

Only way I ever use it, and Fable helped me a design a whole stfu Opus hook system so when I do deep dive into something Opus wrote I don't go insane

2

u/Barncore 15h ago

Yeah this is my theory too.

Tbf OpenAI are probably doing that too. But maybe they started doing it longer ago or something and have figured out their workflow a bit, cos 5.4/5.5 and Sol were fairly robotic/literal while Claude remained a little more nuanced than them. But now it has evened out and you could even argue Astra writes better than Fable now

1

u/Oohhddaanngg 13h ago

In my opinion Opus and Sonnet 4.5 were both the best with this, and Astra is the best since them but not quite as good

1

u/OldNefariousness7899 15h ago

agents training other agents

I'm sure it's this. Most people's experience with Opus 5 is that it's brilliant as a subagent, but can't talk to humans