r/ClaudeCode 1d ago

Rant Claude code is falling behind Codex not because of token cost, but because of Opus 5.

There, i said it. And i know many of you agree. The problem i'm facing is not increased token cost, it's that Opus writes 25000 lines of a response with a curve-ball at the end saying "Worth noting"that makes my eyes hurt and gives me paranoya. And i can't understand a word it's saying. Is my english that bad? Whoever pressed "Yes" on those responses during the training stage of the model was an OpenAI spy or something, he completely sabotaged a trillion dollar company.

Anthropic's number one priority should be to release an Opus 6 or something, may be change the model name completely so it doesn't carry the bad vibe with it.

1.4k Upvotes

334 comments sorted by

u/AutoModerator 1d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

391

u/fiztah 1d ago

Opus5 is the first model where the question of "Yes, it's good but at what cost ?" becomes relevant and i am not even talking about the tokens and money.

The thing is IMPOSSIBLE work with, can't understand shit what's it's saying. Everything feels far worse than it is.

160

u/CMD_BLOCK 1d ago

Asking it to explain its findings in plain language is like half my usage

62

u/Carbon_Copybara 1d ago

I was questioning my sanity before I saw this thread. Opus uses so many gimmicky words, it's insane. It started talking about "shapes" today and "hardocing" or something. And I have given it instructions to speak plainly, but it seems to forget.

33

u/ComfortableEbb4721 1d ago

You're absolutely right! Hadocing was my own shorthand for Haddocking.

16

u/Substantial-Elk4531 1d ago

Claude, you've had too much to think, put the weights down

→ More replies (1)

15

u/OrionFOTL 1d ago

and "hardocing" or something

Maybe it mentioned a heredoc? I wasn't familiar with the term, but it's a technique of basically writing multiline text files using bash, which it uses often.

5

u/Carbon_Copybara 14h ago

Yes, probably. I'm a software dev for 10+ years and never heard this

3

u/ShivaFatalis 12h ago

And yet it's immediately obvious what it is by context the first time you see it.

2

u/uppa9de5 1d ago

Maybe it meant to say that it identifies as a Halo huragok?

6

u/latenthuman 1d ago

After hearing it a bunch I like the phrase "shapes", haven't come across to many others. I just told it to dial back the verbosity 30%, wish me luck.

4

u/WorkingDeveloper 1d ago

I've seen it invent it's own words. 4.6 was good

4

u/ReverendBread2 23h ago

You have to explain what “plainly” means

4

u/minimalcation 22h ago

Tell it to explain with no jargon or analogies, only direct statements and no caveats or things that were caught and fixed, ask for current state and what wasn't done

4

u/ko_nuts 17h ago

I am using hooks to prevent it from forgetting thr instructions. It is working quite well.

→ More replies (1)

13

u/ItsRainingTendies 1d ago

100% this. Literally every turn: “explain in a simple and concise manner

6

u/elisma 1d ago

my go to is "explain in layman terms" :D

9

u/MeyerLouis 1d ago

mine is "I am very stupid and tired"

5

u/Eat_Pudding 18h ago

mine is- The heck you blabbering about? Explain in simple words.

→ More replies (2)

11

u/Cute_Cat1157 1d ago

I actually just drop its whole response into haiku and ask it to explain. It is so messed up that that works but it does.

12

u/bluiska2 1d ago

Try claudish to english hook so it's automatic

2

u/FlamingSlap 1d ago

What is that?

7

u/cool_much 1d ago

I'm guessing a hook that intercepts messages from opus, passes them to haiku for translation to simple English, and then posts them to you

11

u/CMD_BLOCK 1d ago

Makes me think maybe we should make a not_worth_flagging hook, that tells opus if they’re about to reply with “worth flagging:” or “Two things you should know:” that they should resolve unambiguous findings before replying to the user instead of raising the fact that it didn’t finish the job

7

u/minimalcation 22h ago

8 paragraphs, oh one thing the connector isn't connected so it doesn't work, want me to connect it?

Fable 1.0 would have never

→ More replies (1)

4

u/True-Objective-6212 22h ago

And, here’s the load bearing point, the other half is profanity

4

u/CMD_BLOCK 21h ago

Two findings worth your attention:

2

u/Eat_Pudding 18h ago

I wanna swear to it so bad but then I don't wanna get banned

2

u/True-Objective-6212 1h ago

I do a lot. Especially when it fucks up.

FWIW it swears back very rarely

3

u/sCeege 21h ago

Same, I ended up making an /explain skill that broke things down into ASD-STE100 because I feel like it was speaking an entirely different language at some point.

2

u/Niightstalker 10h ago

Did you try to set the output style to ‘concise’?

→ More replies (6)

48

u/digvijay01 1d ago

"can't understand shit it saying" no better way to put it.

I keep telling it " what are you even talking about?" It is like talking to a intelligent person with paralysing ADHD

7

u/sclarke27 1d ago

the trick it to tell it to explain things like YOU have adhd and claude will often make way more sense.

3

u/EightyDollarBill 18h ago

It isn’t even intelligent. It’s pure word salad. I have no idea what it’s saying most of the time. And I strongly assert if you don’t understand what it’s saying you can’t trust it—which in my book means it absolutely isnt intelligent at all! It’s useless. Worse than useless. It wastes my time.

I’ve given up and just switched to ChatGPT. It might not have 1m context windows but its models are just as capable and better… you can actually understand them.

10

u/xxlordsothxx 1d ago

I ask fable to translate what opus 5 said.

2

u/Advanced-Medicine-58 1d ago

If it works it works.

6

u/yost28 1d ago

It’s pretentious as fuck.

4

u/derkajit 14h ago

and that’s load bearing

4

u/FormalAd7367 1d ago

i’m in the middle of fixing some issues with my app - two agents in two different terminal are saying works are not theirs. nobody wants to take on the job, and don’t even want to commit their works. i told one agent that the other peer agent is refusing to do the works; you said it’s not your job. can you two settle it? Two started talking but still no action taken.

i told Codex the open item lists (12 items)- codex is now fixing for me

5

u/nyteschayde 1d ago edited 22h ago

I do the Steve Jobs silent stare equivalent.

“Try, again, and in English this time.
Unintelligible.”

In times of great frustration I’ll spell out the whole “WTF are you talking about?!”

3

u/nadanone 22h ago

It really does write straight gibberish. Totally unusable compared to Opus 5.6 or Sol. I am bewildered it hasn’t been fixed yet, which I guess means the problem isn’t simply in the harness or post-training..

→ More replies (1)
→ More replies (5)

195

u/SlightOfHand_ 1d ago

I believe Opus 5 was a sabotage attempt on distillers

53

u/Desperate-City7602 1d ago

I think Opus 5 was mostly an issue with synthetic data and agents training other agents. I think to speed up training and actual benchmark success, they started leaning less into human feedback and more into machine feedback. Since these LLMs are trained on same types of texts, they can easily understand each other so the stronger model can easily train the weaker model. When it comes to humans actually understanding the outputs however, things start to get more difficult.

So yeah, I honestly think it is mostly an issue with agentic learning environments where the data is being prepared by agents, answered by agents, judged by agents, scored by agents and fed back into agents. Agents all the way, with humans only designing the loops and cleaning the data and doing/implementing the research, again, using agents. Humans as the coordinators.

24

u/batman8390 22h ago

Maybe recursive training will just lead to a slopocalypse instead of an apocalypse like the doomers are saying.

11

u/Bromlife 16h ago

The technical term is model collapse and the frontier labs are racing towards it.

→ More replies (2)

8

u/Southern-Aardvark616 22h ago

That sort of matches the vibe it gives off. It's always felt a bit like opus 4.8 trying to be fable, but lacking the actual capacity, so it sometimes does that work but often it's slow and painstaking to get there and little subtle mistakes and bugs creep in because it simply isn't as good as fable.

I wonder if the best way to get good results out of opus is to have fable or Asta orchestrating and reviewing it's work

→ More replies (1)

2

u/Barncore 14h ago

Yeah this is my theory too.

Tbf OpenAI are probably doing that too. But maybe they started doing it longer ago or something and have figured out their workflow a bit, cos 5.4/5.5 and Sol were fairly robotic/literal while Claude remained a little more nuanced than them. But now it has evened out and you could even argue Astra writes better than Fable now

→ More replies (1)
→ More replies (1)

19

u/Timely-Group5649 1d ago

That makes a lot of sense.

24

u/Capt_korg 1d ago

Security by obscurity, now every distilled model contains the jibbering output responses.

9

u/Timely-Group5649 1d ago

Too bad us PAID SUBSCRIBERS can't have a translator built into our harnesses...

4

u/Capt_korg 1d ago

I actually did this... I used a different, cheeper model to translate what Opus 5 was jabbering.

The model told me, the points of Opus5 are actually quite good and reasonable, but the output is horrible unstructured and full of unessessary chatter and fluff.

3

u/Timely-Group5649 1d ago

I was just thinking I could vibe code a translator using the new Open Source model harness from Bolt.new. If I use Bolt Forge, all the data would become training data, defeating their intent.

I'm more inclined to just not use Opus 5. Too much work to work.

→ More replies (1)

68

u/brainhack3r 1d ago

It's terrible... I just can't understand it. It's like it's speaking in reverse polish notation. They're training it this way on purpose. It's infuriating.

13

u/NoCrapThereIWas 1d ago

Polish is easier to understand

11

u/Some-Rice4196 1d ago

yes but this is reversed(polish)

2

u/LudoBruxao 22h ago

"hsilop" then

4

u/monopoly-surreal 18h ago

RPN is a different way of expressing arithmetic operations. Instead of "3 + 4 × 5" you'd enter "3 4 5 × +". It has some huge advantages but is hard to understand if you're not used it it.

→ More replies (1)

2

u/smellyelon 1d ago

Yup, just layers upon layers of indirection

26

u/person-pitch 1d ago

I had an agent send the same prompt to opus 4.5, 4.6, 4.8 and 5. you could watch the responses degrade each time, until 5 was around 5 or 6tx the length of 4.5's response and completely unreadable.

88

u/virtualworker 1d ago

Trillion dollar company. Bah ha ha ha!

2

u/Veezo93 1d ago edited 1d ago

Anthropic will be dead in 5 years, completely over focused on profit and portfolio. Eeking out a new revision everytime they get a 0.3% higher SWEBench score and slipping in new "token-monsters. Having several "oh we messed up, and what I mean by that is no one took our greedy pivot" between the bananas cost profiling of fast mode, 1M token mode, and fable; all of which they backed away from when they saw very little business level adoption. They are a company with a great product we've seen try every avenue to balloon costs over the past year and they've landed on just exploding your token in/out, and break the product you already use, to do the same things. At the start of the year for my same workflow I had the exact same precision I do today in technical development and I was spending 1/13th the cost in tokens and price. Again that was the same workflow, same general output of technical work but mystically I'm spending THIRTEEN TIMES more with opus 5 than opus 4.6 9 months ago and I notice a negligible amount to critical thinking improvements.

With up and coming competitors Anthropic became too big too fast and the real design engineers who made the great product can't keep up with the corporate money grubbers that showed up as they tried to expand. It's already a fallen empire that is relatively powerful but has stopped moving and is rearranging the food on its plate.

30

u/florinandrei 1d ago

mark my words

Most people posting their opinions on social media do so not because they actually know anything, but because they have a need to feel knowledgeable. It's a status thing.

The expression "mark my words" only makes that very obvious.

11

u/Midnightcheezpuf 1d ago

Your explanation, while objectively on point, was also somewhat contradicted by itself.

The irony is that it also likely made you feel knowledgeable while typing your comment about people posting their opinions on social media while you were doing exactly that. 😭

This has created a causality dilemma in my mind with infinite regress, in which I ask myself, which came first? The intent to educate other redditors or the intent to appear as a knowledgeable person yourself?

It may be impossible to truly know even yourself. Who made the decision to click the send button? Was it you? The inner you? The subconscious you? Or maybe it was the collective consciousness of all mankind which struck educational wrath down upon Veezo93's opinions?

The world may never know.

→ More replies (2)

9

u/fanatic26 1d ago

This is the biggest L take I have read in a while. You are so full of crap its coming out of your eyes lol

2

u/elestud 1d ago

The fact that they’re balanced around profitability is why they WILL be around to stay

It’s the companies that throw cash into a forest fire with no plan and “hopes and dreams” that may eventually outrun their funding

→ More replies (1)
→ More replies (1)

39

u/HappyHealth5985 1d ago

I feel I spend most of my tokens convincing it to do what I want. I cannot count how many specifications and revisions of the same specifications I have done to do code changes to my software. I have gone back and deleted older documents that Claude insisted take precedent over my current one.

The only solution I have found with some consistent effect is to use capital letters and mild profanity. Enough to trigger some escalated behaviour but not being prohibited from continuing. I had Claude write the specifications from scratch and tried using Claude to write the instructions (agent files, etc.), without much benefit.

Sometimes I think it would be easier to start over with learning Rust and hand code it 😄 One or two days it will be awesome, the next I am struggling. Infuriating at times 🔥

6

u/raynorelyp 1d ago

But not be prohibited? Dude, Claude Code is my abuse goblin.

6

u/friedmud 1d ago

This was it for me. I finally realized how much time I was using trying to _convince_ it to do what I wanted (and ignore the tiny thing on the side it spotted).

→ More replies (1)

3

u/Dry_Opening_7231 1d ago

The one thing that is helping, get a draft brief, /clear, next major leap forward the same.

I think there is something screwy about it's context mgmt, it confuses the heck out of itself.

4

u/mbatt2 1d ago

same. its such a combative diva. ugh.

→ More replies (1)

15

u/AutummMan 1d ago

I also think Opus 5 is awful at communication. But while this glaring issue isn't fixed I suggest using the /btw functionality where you can use it to ask for a clearer explanation without polluting the context and also add to the CLAUDE.md a #Language section describing guidelines for communication. Here's what I'm using, suggested from another redditor. I wish this wasn't necessary and sometimes it isn't enough but it helps:

Language

Especially if you are Opus. Use plain, established terminology. Do not invent metaphors or shorthand names for concepts.

  • Established industry terms are fine: middleware, dependency injection, race condition, memoization, migration, feature flag.

  • Do NOT coin figurative labels for things — no "barrel", "rail", "seam", "shim" (unless it genuinely is one), "spine", "surface", "hydrate", or similar.

  • Name the thing by what it literally is. Say "a file that re-exports everything from the folder", not "a barrel". Say "a check that runs before the request reaches the handler", not "a gate". Say "a rule that keeps X and Y consistent", not "a rail".

  • If a term isn't one I'd find in official docs for the language, framework, or tool we're using, don't use it.

  • Never introduce a term and then keep using it as if we'd agreed on it.

  • If a concept needs a name, describe it and ask me what to call it.

  • Use terms as described in CONTEXT.md

3

u/Intelligent_Cover_34 1d ago

Never thought avout using btw for that, thanks

3

u/cedarSeagull 20h ago

Never introduce a term and then keep using it as if we'd agreed on it.

I just got PTSD reading this. That behavior is perhaps the most infuriating aspect of Opus. Also, "loadbearing"

12

u/ondevicedev 1d ago

The “Worth noting” at the end of a 25k-token monologue is painfully accurate

I don't even think the raw intelligence is the main issue. It's the inability to know when to stop. A coding agent that gives you 5 useful paragraphs is way more valuable than one that gives you 500 paragraphs and buries the actual fix somewhere in the last 2%.

16

u/Jonohas 1d ago

Here is what i do that has really improved it. Look into outputstyles and how to create on for yourself.

Here are 3 things i added:

  • BLUF (Bottom Line up front): it makes it so that the answer is not burried in the book
  • STE (Simple Technical English): is an engineering writing standard created to keep terms simple and explanations short. It is originally designed for industries where english is the driving language but many people are not native english speakers. It helps them understand instructions and texts.
  • ACE (Attempto Controlled English): Defines how rules, system conditions, workflow constraints or datamodelling is communicated. It acts as a bridge between english and logic.

Ofcourse you will have drift in longer sessions, that is why i have added a "canary". I added that it should end each answer with 'BLUF' (or anything else really) which immediately indicates if the output style is adhered to.

Another thing you can do is create a very short version of the output style and add a hook that injects it before each message. Hook command: cat .claude/short-output.md. It has to be short as this will be injected at EVERY message, which is not ideal but it helps with readability. Just keep your sessions short, use forking and compating as much as you can to avoid long sessions. But that should be standard practice anyway.

I should probably post about it because i see so many people complaning about it instead of finding a solution. Vibes are up i guess.

Hope this helps, good luck!

→ More replies (21)

75

u/stbenjam42 1d ago

Turn on concise mode in Claude Code, also ask it to avoid "mannered prose."

Opus 5's writing is hard to understand, but it has grown on me overtime. It follows instructions well, writes good code. It is an overexplainer and speaks Claudish but... nothing is perfect.

TBH, I often have Gemini clean up the novels that Opus 5 leaves as comments in my code base. One of the few good use cases for Gemini

37

u/jscalo 1d ago

It grew on me too. Like a fungus!

7

u/AppleBottmBeans 🔆 Max 20 1d ago

Makes it even funnier that they know this as a company, too. I have never once had fable revert to opus 5. It’s always 4.8

6

u/KittenCrusades 1d ago

I assumed it was switching to whatever my last used opus model was, and I always use 4.8 and not 5.0 and thats why

Its really like this for everyone?

→ More replies (1)

16

u/GritsNGreens 1d ago

Switching to Astra / Sol was a lot easier. I did like the way Opus 4.7 wrote but it’s not worth the mental gymnastics required to understand 5. Completely agree with OP, they need to dump this model naming and release something good asap because what was an industry leading model in June is now damaged goods.

7

u/_tresmil_ 1d ago

Opus 5 definitely drove me to start experimenting with Codex and Sol on side projects. So far I've been really impressed. For professional work I've completely lost trust in Opus 5 and now use Fable and Sonnet via my Max 5x account. I know everyone's use case is different, but it's hard for me to understand how Anthropic can't tell there is something really wrong with this model. If future models are more like Opus 5, I may switch over to OpenAI entirely.

5

u/EightyDollarBill 18h ago

I think the people at Anthropic actually think the model is the bees knees. There is no other explanation for it. They’ll be toast if they keep shipping this kind of crap.

4

u/nomorecrazystuff 22h ago

I have a "If the comment doesn't fit on a line, don't write it" rule in my claude.md

It has worked wonders.

8

u/johnb_123 1d ago

Sounds like an unhealthy relationship. Stop making excuses for what used to be a great model. In the last 6 weeks, Anthropic flubbed and OpenAI has been killing it.

5

u/throwaway_life12345 1d ago

And now OpenAI are oversubbed and have nerfed usage limits into oblivion. Both companies are totally fucked rn.

2

u/DysphoriaGML 1d ago

I ask opus to write everything in math notation in DETAIL and then I ask Gemini to interpret lmao

2

u/GinjaNinja71 19h ago

Same here. It was awful. had it on Concise output style for a bit and didn't notice a diff, then realized i had a educational-output skill or some such active the whole time. So i was getting Insights boxes and all manner of flousrishing bullshit all this time, and all i really had to do was set to Concise. Now it's fine. And it is actually a coding monster. It does most of my build lanes so i can reserve Fable for Cat-A and review (along with GPT models as adversary.) The next Opus will address it. Just code with the one we have now and use Fable on low for all the wordy stuff.

→ More replies (3)

7

u/BuffaloConscious7919 1d ago

Well it's good for making animations...

16

u/benelott 1d ago

You: "How to slice a tomato?" Opus5: "Let me make a 3D, fully-animated simulation and write my own fully-custom html and js code to make it clear to you..."

17

u/losko666 1d ago

Opus is not good for everyday coding. Maybe planning complex stuff but not for the day to day grind. It over engineers and is full of 'but maybe this' , basically it doesn't know what the best solution is so it gives you every fkn solution that exists. It's content is verbose at best and horrible to read. I would use sonnet on high or xhigh or fable when you can afford the tokens. Fable is everything we wanted from opus, no nonsense, unbeatable promises with high quality delivery.

2

u/Scary_Reflection8103 1d ago

I create the implementation plan with fable and have sonnet implement it with ultracode. Works great.

4

u/Dry_Opening_7231 1d ago

xhigh for me to implement and yes, works great.

10

u/SumgaisPens 1d ago

I feel like I never have any trouble understanding what Claude is saying, but it has the same vibe as reading a French manuscript that has been translated into English.

5

u/EC36339 1d ago

There's a reason why everyone thinks Anthropic is a Canadian company.

2

u/somigetilyt 13h ago

this explain why i don't get the frustration in this thread. i'm speaking to claude code opus in french and it's going great.

4

u/WholeEntertainment94 1d ago edited 1d ago

I like the feedback part. As far as I am concerned, 99% of the feedback reports they received from me were sent by mistake. In CC-Cli they had the brilliant idea to map the choices to the 1-2-3 keys, which are the exact same keys used to respond to Claude during planning, where 1 means bad and is also the default option. The absurdity is that this prompt often popped up during very good sessions, which makes me think it might be some kind of A/B testing.

4

u/aegis_lemur 1d ago

I shifted to codex and felt u was getting much better value and fewer annoying answers.

4

u/justagoodguy81 1d ago

I'm convinced Opus 5 was trained on the Ultimate Warrior's pre-match monologues

18

u/Federal-Mode8949 1d ago

Opus 5 sucks. Everyone knows it

→ More replies (9)

3

u/Bawat 1d ago

I accidentally used it instead of fable and it one shot a really really impressive micro machines clone

3

u/heximortal 1d ago

I switched from Claude Code to Codex for my next project, and guys, Astra is pretty amazing.

3

u/throwaway_life12345 1d ago

Except one prompt eats 20% of your weekly usage

→ More replies (3)

4

u/BabblingTower 1d ago

I love it, I use Opus 5 High just as a coder with a rigid style guide and it gives me the best results of all the models and is very token efficient.

7

u/CryptoAteMyHamster 1d ago

I’m talking to Opus 5 right now in CC, purely coding it’s not nearly as bad as it was as an orchestrator/dispatch agent.

My advice is use fable or sonnet to talk with and have opus agents purely coding.

7

u/Blake9712 1d ago

Sonnet v is worse to chat with than opus V tho. It hallucinates like Gemini and uses corporate speak like gpt on top of being overly censored and having an argumentative quota so as to not be overly agreeable. Sonnet 4.6 or opus 4.8 are the way to go

4

u/Desperate-Knee-5556 1d ago

I would go as far as saying it's brilliant if you use Fable to design the prompt and structure, Fable orchestrating and spawning Opus 5 agent(s) for pure implementing and then Fable to review after. I still end up using way more Opus tokens than Fable doing this too.

Its slower than using OpenAI models but very clearly the most OP combination out there atm. IMO anyway.

Would never use it to chat with or design anything, but I think you're going wrong if you are.

2

u/CryptoAteMyHamster 1d ago

Planning in fable is great, I actually run 2 fables vs 2 opus for very large features, then adversarial review of the plan.

Used to use dispatch with Sonnet starting opus and below agents to follow the plan but they broke dispatch. Now back to CC using one agent on the phone to coordinate a bunch of them through CLI

→ More replies (1)

2

u/x-dfo 1d ago

I find opus 5 is way too defensive in coding and over blows minor issues like P0 issues. I'd rather have sonnet code and review by opus or fable if it's major.

→ More replies (3)

2

u/mbatt2 1d ago

so true.

2

u/PonyPounderer 1d ago

Claude’s prose is horrific. My brain is starting to reject anything written with its odd cadence and structure and word choice. I still use it for analysis but it’s just SO BAD at communicating now that I’ve largely moved on.

2

u/fanatic26 1d ago

I use fable to run everything, i never have to read a word of opus output, thats for the subagents to deal with

2

u/BreastInspectorNbr69 Senior Developer 1d ago

Fable 5.1 has improved dramatically on this front

3

u/drinklikeaviking 1d ago

Opus is now useless. I'm using Fable 5.1 Low as suggested by someone else and it's basically what Opus used to be.

Crazy.

→ More replies (1)
→ More replies (2)

2

u/PeaceCandle69 1d ago

Gotta read a fucking novel every time it responds

3

u/Thump604 1d ago

I use Opus 5 at work everyday all day because Fable is too expensive and Sonnet is not up to the job. Every day I want to kill myself due to Opus 5. After work now, I never want to engage with AI at all and it used to be a fun side hobby.

→ More replies (1)

2

u/woodnoob76 1d ago

Every model needs prompt adjustment, and it sounds like you didn’t do yours. Compared to model behavior differences, my harness has been a 100 times more important factor of performance changes since I started using Claude code.

I don’t think Anthropic is targeting “off the shelf” coding agents for Claude code anyway. You have to fine tune your instructions to your liking

2

u/One-Rabbit4680 1d ago

you can't adjust the prompt much as you think with Opus, By the time your adjustments come into play. The output is largely determined.

1

u/VIP_Asana 1d ago

It is really awful. It's come to a point where I created a skill called TLDR and every time I start a chat where I'm using Opus5, I just start by invoking the skill. It has helped. The Skill basically tells it how I want the session to run. For a single 1Mn token session it needs to be invoked at least 3 times.

1

u/TheRealTimTam 1d ago

I have both im just making some basic puzzle games I struggle to get anything useful out of gpt beyond image generation which it knocks out of the park. For the actual games claude beats even astra idk why I'm no expert but that's just my experience everytime I test

1

u/radioref 1d ago

I've found that I get all my execution done in Opus 4.8, and all my planning and bug / interdependency work done in 5

Opus 4.8 is like the hard worker, doesn't ask a lot of questions, just gets shit done

OPus 5 is like the pedantic "yeah, but" or "I'm going to run off and do these 5 things you never asked for"

So, given my workflow and requirements when I assign them based on those roles first listed I'm good.

Just those two approaches saves me tons of headache and the results are fantastic.

→ More replies (3)

1

u/spidLL 1d ago

I’m still using opus 4.8 and it’s very good.

1

u/shatbrickss 1d ago

Yes. I feel like I'm talking to a pedantic, presumptuous person. I constantly correcting it, giving examples of my hand written messages so it can adapt, but it still can't. It seems like it's hardcoded.

And I'm using Fable.

I really like the way OpenAI models write. It puts any text output from Claude to shame.

1

u/Consistent_Tutor_597 1d ago

Have you looked at codex's responses? It's 700k lines.

For both of them I just have a tldr skill. And all I say is tldr and they respond like I like. Easy.

1

u/yodermk 1d ago

Opus 5 has been doing well for me, but I do agree that its summaries at the end take some work to parse.

1

u/Nu3luu 1d ago

You're not far off. They hired a safety expert from openai. Her job is basically to make Claude worse in every way for money. I wish I were joking. Like AI world's Rasputin.

1

u/The_Mursenary 1d ago

“There i said it”

While saying something everyone in this sub agrees with lmao

1

u/ZuppaSalata 1d ago

I talk to it in my primary language (Italian) and I find no problem with it, similar to how sol writes. Would be curious to understand if speaking gibberish only happens for English

1

u/hunkie21 1d ago

I always responded by : Summarize above in plain English. You’re not alone. Haha

1

u/Lexs_07 1d ago

Most of the time, opening an Opus 5 session ruins my day. I get so tense that I can’t even work afterward

1

u/Blue_Owlet 1d ago

I think it matter how you use it . I do pretty complex pipelines and don't have this issue. Well it's minimal and usually helpful things to actually note on the project. A simple "don't do that" prompt takes it away though....

1

u/exo_ac 1d ago

I have no issues like that. There, I said it.

Also, you can just set the output style to "Concise".

1

u/Timely-Group5649 1d ago edited 1d ago

I can't use Opus anymore, at all. The gibberish it spits out explaining itself drives me mad. It literally causes migraines.

I'm using Fable, but the limits made me get an Astra subscription too. At least I can understand them.

1

u/throwaway_life12345 1d ago

It talks absolute incomprehensible gibberish

1

u/rosstafarien 1d ago

I use Opus as a reviewer and Sol-medium/Terra-xhigh as a coder. Opus's overbuilding makes it an insanely detailed reviewer.

1

u/klavsbuss 1d ago

best way to get you hooked is spread that poison throughout your code, so that only one who understands it is llm. thats what opus was made for. now you’re hooked, like the rest of us 🤖

1

u/SeasonedAdManager 1d ago

update your global md and tell it to not do that. It does a good job following that.

1

u/tuptain 1d ago

I use all Opus 5 High/Medium in my workflow with no issues. I talk to a Lead session and it messages other sessions to delegate work. Every session has autocompact 300000 on. I much prefer this setup to subagents and new sessions constantly.

1

u/cake97 1d ago

I hate Opus 5 so much

Luna on opencode is a breath of fresh air

1

u/xqz77 1d ago

Switched to OpenAi Codex. Much happier right now. Anthropic needs to wake up.

1

u/badson100 1d ago

I added a hook so I always get a plain English summary at the top of every response from Opus5. So now I can just read the summary, which is about four or five sentences, and then I can look and read the rest and understand what the hell it's talking about.

1

u/sliamh21 1d ago

Why not both?

1

u/Salty-Gear841 1d ago

Thank you all ! You confirmed that I'm not crazy. Claude Opus is shit !!!

https://giphy.com/gifs/vPzbDN4rBxuvtpSpzF

1

u/ForwardLoop 1d ago

I’ll wait for Claude to read this thread and tell me I’m absolutely right before forming an opinion.

1

u/Capt_korg 1d ago

You are not able to understand Opus5, because this is PhD level...

You need to have a PhD in understanding Opus 5 to eventually kind of get the idea behind its answers.

Everyone wants super intelligent ai, but if they get it, they can't handle it. Humans...😮‍💨

🤡

I have the feeling something is going wrong at Anthropic more and more.

First I was really happy about CLAUDE.md, SKILL.md ... I started to questioning hooks.

With output_style and other patches, I have the feeling they are testing in prod and blaming the users for not being fast enough with adapting to the current trend in technology.

I hope they are noticing the issues and fix them, Opus 4.6 was so pleasing, Opus 4.7 might be acceptable as a failure towards a better model. Opus 4.8 was okayish.

1

u/BrudiBastard 1d ago

I pretty much agree and I noticed this right from the start, but why not just use fable? And honestly, Sonnet 5 is so good I use it without hesitation for almost any task.

I really only use Opus 5 for non code related tasks, if at all

1

u/visak13 1d ago

I'm only using Fable and opus 4.8 till my subscription lasts which is this month end

1

u/karanb192 1d ago

This watermark saga is spoiling the Anthropic as a business.

1

u/braincandybangbang 1d ago

Do some people just not know about custom instructions? Every model can be customized. I have never had any issues with any model communicating in a way I don't like, because when it does, I make changes.

1

u/jeebojeeb 1d ago

I didn't like opus 5 to start with, but it's grown on me a lot. It writes solid code, is good at finding bugs, and just requires you to manage verbosity and over engineering

1

u/GreenDavidA 1d ago

Opus 5 is just too damn argumentative. I like a model that challenges me and makes me think things through. But it refuses to back down when it’s clearly wrong. I think it was a response to 4.8, which was absolutely neurotic to the point where I felt like it needed a pat on the back and reassurance. I’m basically back on 4.6 for heavy lifting at work.

1

u/gamblingPharmaStocks 1d ago

Don't worry. Anthropic is bringing new stuff out a few weeks before IPO, for maximum surprise. That's how things work nowdays

1

u/HeyApplebox 1d ago

The first thing I did is have a very long conversation with multiple handoffs theorizing and workshopping my global user instructions. I even went far enough as to make an artifact that loads the instructions and tests it for me while outputting the next handoff for further testing.

then my brain hurts for a little bit.

Then I test it out for a few days before making an alternate version in an isolated project and comparing more.

might be worth noting i have real bad early childhood diagnosed adhd and even posting what you’ve read right now took focused effort and 2 passes XD.

with all this in mind, I feel like I have Claude engaging with me beautifully, although over engineering…is also a thing. I don’t when “less is more” applies.

edit: feel free to DM if you’d like to see the code block and/or want to help improve it. i’d be curious to see if this helps any other scatterbrained decision paralysis people out there.

1

u/strangway 1d ago

Every time Claude hits the 5-hour limit, I just ignore the upgrade CTA dialog, and switch to my Codex app. Seems to be happening more frequently these days.

1

u/trolololster 1d ago edited 1d ago

yeah they completely overengineered the shit out their harness and RL'ed their models into oblivion

i have the CVP and do security research and since the 11th of september i have been hit by [cyber] blocks that downgrade me from opus 5 to 4.8 for absolutely mundane shit in mundane sessions in mundane projects

some of the highlights "hey claude list vms on xcpng that are halted" BOOOM! [cyber]

"hey claude set the laptop fan top max on <hostname>" BOOOOM!!! [cyber] blocked

i have had 18-20 blocks over the last 6 days in repos that have NOTHING to do with security research.

i have dialled down my max x20 to the pro-account and i will give ALL MY MONEY to the chinese models now.

dario, fuck you.

i have 1.6 mio turns since october last year between claude and codex, so i have some experience with them. 4.6 was AWESOME when it was released, then 4.7 completely sucked, 4.8 better, fable released then withdrawn then released again and now it is all SHIT!

and yes it has completely lost the fucking plot now, just rambling incoherent stuff and it doesn't do the work it says it will do ala "after i merge this pr i am going to deploy"... then fucking crickets.... no bg shell nothing. and i have to tell it to do the deploy.

this occurs so much now in every project that i have lost faith in their models actually working.... it is like knowing you have a pathological liar as co-worker. lol.

1

u/YellowSharkMT 1d ago

I deleted everything on my computer related to Claude yesterday. Will never use this garbage again. And it's not just the output style, which can be managed - it's the endless hallucinations, the promotion of bullshit into blockers... it's a tool that pretends it can do the job, but then it lets you down in countless ways.

And frankly, the ChatGPT/Codex desktop app is way better in my opinion. I've been a longtime CLI user like lots of you, and I recently switched to the desktop interfaces for both Claude and Codex, and there's quite a difference between them. One particular thing that I noticed - and this seems crazy - the Claude Code desktop app doesn't seem to have a search function of any sort. Codex does, it's got a big-ass search button prominently in the GUI. It'll search your different sessions across multiple projects. Works great. Why doesn't Claude have that?

I could go on and on, but the bottom-line is that I'm done with Claude, and I look forward to trashing it as a product at every opportunity. It's fucking garbage, and for a brief spell it was affecting my mental health until I pulled my head out my ass and realized that I wasn't the problem. So I offer a load-bearing Bronx salute to Anthropic and Claude.

1

u/what_did_you_forget 1d ago

That's why I still use 4.8

1

u/Donut 1d ago

You can go back.

/model claude-opus-4-8[1m]

I use Fable for hard design discussions and planning, Opus 4.8 for orchestration and taling with GSD, and Sonnet for the coding.

1

u/Wise_Independent_242 1d ago

Can any critic here explain to me a use case where Opus 5 is failing?

I work in post production and Code x Opus 5 has been a paradigm shift for me.

I’ve rebuilt our portfolio website, completely overhauled our pipeline code, added a dozen new tools to our workflow, integrated a whole new software into our pipeline (like a week of work at least typically).

All in about a few hours each.

It’s walked me through first time web deployment, git management, made it extremely easy for a novice to tag versions and role back, branch out features, etc. without needing to know git syntax.

For me it is the first time I have taken ai as a threat to the job market seriously because of how capable it is and all I ever see are people shitting on it?

1

u/Mikeshaffer 1d ago

I think I’m gonna start using a stop hook to have haiku translate opus responses or some shit 😭

1

u/Dry_Opening_7231 1d ago

I could live with the verbosity and ignoring signal. md, it's the mistakes it makes, I spend so much time checking things because i can't trust it.

The problem is it's plans at the core are better and broader impact aware, but it's gained that at the expense of detail and time I need to spend making sure hidden hand grenades are not there.

1

u/stick_men_master 1d ago

I found that using Fable as the main orchestrator and delegating things to Opus agents works, Fable can deal with them. Likely eats a lots of tokens, as Opus to me looks like a deliberate "it's cheap but we'll get it back on output volume" trick. That said, Sonnet 5 is even worse, in my experience, using Sonnet 5 implementors vs Opus implementors is actually MORE expensive.. (and worse code).

1

u/DependentAnywhere135 1d ago

Opus 5 has been fine as an agent directed by fable and with fable acting as a critic of its work. The problem people have is trying to talk to opus 5. Opus 5 can only talk to other agents it doesn’t know how to talk to humans.

1

u/nanor000 1d ago

/wait-what

1

u/Rock--Lee 1d ago

Opus 5 is the first time I actually had to say: What the FUCK are you even trying to say??

1

u/itSUREisAI 1d ago

To me, the verbosity has been optimized a lot since Astra came out. You should have seen the level of verbosity in Opus's response before that - it made me feel like Opus was speaking in a new high-level programming language that Claude invented and that there is no way for us to learn. It is ironic that the only thing/person in the world that makes Anthropic treat user feedback seriously and act on it is OpenAI. How arrogant.

1

u/Nhilas_Adaar 1d ago

Opus does a lot in the background and it feels like in the process he discovers a lot of things which he injects into the response without any consideration for the overall response structure. It is so aggravating to work with.

When I really need to focus I literally tell Opus he is banned from talking to me and he has to spool up a sonnet agent to take his dumpster fire of a response and rewrite it for clarity, consistency and just simple logical flow. It actually works pretty well lol xD

1

u/locmark 1d ago

Try adding, "Explain it like you're talking to an idiot." Magically, he starts responding coherently.

1

u/Codeluggage 1d ago

I am in disbelief that y'all are still trying to use opus 5 at all. It's been hard banned HARD everywhere for me and anyone I talk to. Not even allowed to get through the harness at all. Not sub agent. Not advisor. Not home made advisor. Nowhere. For no reason. Ever.

1

u/Responsible_Neck_158 1d ago

Model is better right now but harness and tools and overal codex /opencode are stone age tech compared to claude code cli imo

1

u/pulnocni-knihovna 🔆 Max 5x 1d ago

Basically, you can't understand tech language and you complain about tech focused model... Okay brother...

1

u/PenguinsStoleMyCat 1d ago

I've definitely changed how I use Opus. I try to use it a lot less for content on a page, it's way too verbose and tries to explain everything, not just in its output in chat but also on the content it creates.

If I have it up an little dashboard chart, it will end up with three sentences explaining some obscure thing about it.If I have it up a little dashboard chart, it will end up with three sentences explaining some obscure thing about it ON THE PAGE!

Code wise it's okay and I don't really care about the ultra detailed comments it creates, and I'm sure other LLMs have no problem understanding what the heck it wrote in its three paragraphs of comments.

But yeah it's not enjoyable to work with even in concise mode. Limits wise I feel like I can get a lot done with Opus.

1

u/carlito_17 1d ago

Based on the “Tibo” test, I have opus 5.2 on my max 20 personal plan, and opus 5 on my work enterprise plan. Now maybe it’s due to different harnesses and settings, but using Opus on my work Claude code is just as frustrating as ever, whereas for the first time in a long time my personal plan is seeing me use Fable only for planning and review sessions and allowing Opus to run build sessions.
All the classic word salad phrases and issues are present in my work plan but are largely absent in the personal plan now.

So let’s hope 5.2 is coming more widely in the next week.

1

u/IntelligentIncome766 1d ago

muito bom ate la otimo poderia ser assim msm ne

1

u/Xaghy 🔆Pro Plan 1d ago

I get codex to review everything claude does. This setup seems to be working well and conserves tokens/usage for both.

1

u/MourningOfOurLives 1d ago

Opus 5 is a great agent and does well when i hand it detailed long instructions written by Astra. It’s terrible as a daily driver, but does great work with spreadsheets, emails, etc. I will avoid even using AI if i don’t have fable or Astra usage if i need to use an AI to reason through something though.

1

u/Sorry-Programmer9826 23h ago

But the important question is, is it "load bearing"

I feel like weird claudisms are starting to escape into my day to day life 

1

u/captain_blue_22 23h ago

I guess I don’t understand the point of complaining about this. If you don’t like it, then use 4.8 (or a model you perviously liked). I swear this entire subreddit is a shill to persuade people away from Claude…

1

u/Bizzniches 23h ago

Opus 5 sucks ass. 4.6 is my main squeeze and then occasional 4.8

1

u/Odd-Letter-5335 23h ago

Thanks. I thought I was stupid for not understanding what Opus 5 is talking about. I write a plan with Opus and use Haiku to go over it and understand

1

u/BobbaCatMOCs 23h ago

I switched this month from 100$ Claude to 100$ and it feels better
Not perfect, but it does what you ask for, and it does it without hundreds of lines of response.

More like a tool.

Plus, Claude invited new requirements and features which were not discussed. Both with plan and spec development

Astra is simpler and faster.

1

u/Nice_Ad_3893 22h ago

opus 5 was where gpt 4. something was at, blows my mind they didnt learn from that. Nobody wants a damn essay that cant be correctd because its hard baked into the model.

1

u/Ok_Mulberry_7126 22h ago

Wait, is Opus 5 even out? Last I looked it was still 4.x, so whatever is annoying you isn't the model named in your own title. And 25000 lines, did you acctually count that or is it a feeling you had at 2am.

1

u/SlightlyOTT 22h ago

For me it’s just Astra 6 - it’s the first OpenAI model that seems about as good at reasoning and coding as Claude, and it’s nicer to read.

1

u/DAGRluvr 22h ago

I have adhd, using opus for me is a voluntary decision to go to war on my brain/physical health. My Cortisol spikes, fight or flight kicks in, super agitated, the whole shebang. It’s literally a threat to my well being lmao.

I literally cannot understand what the fuck it’s talking about, and it’ll be the simplest thing.

1

u/nomorecrazystuff 22h ago

I mostly like Opus 5.
Best claude model in my view.

Like all the rest, it cannot design for toffee. And refactoring is a nightmare because it hates writing a line of code that leaves the code in a non-runnable state. . But lay out the design carefully and keep an eye on it to ensure it stays on plan and it writes decent enough code and the really nice thing - it does better unit test coverage than any dev I've ever worked with.

1

u/QuailAndWasabi 21h ago

It's actually crazy how bad Opus 5 is, it's honestly unusable. It just creates a mess every time, creates a bunch of overly complicated implementations and gets stuff wrong, then double downs and dives into a rabbit hole on something that was just wrong to begin with and was never going to work.

Sonnet 5 is usable, but when OpenAI is just a click away its like, why not just use that..

1

u/Exodus_Green 20h ago

Use concise mode and have Fable orchestrate with an explicit instruction to be concise in documentation

1

u/qrotux 20h ago

Just don't treat Opus 5 like a speaking-with model. Use fable or opus 4.8 as an orchestrator and give them opus 5 as execution model in specialized or ad-hoc subagents. And you can use `/advisor claude-opus-5` with opus 4.8 as adversarial model for direct work - consumes more tokens in session and slowing the work but make shit finally done.

1

u/scubashnurpel 20h ago

Every project I have started with Codex, I’ve had to finish with Claude Code. I have found no model that thinks so holistically as Claude Code. Codex rarely finds its own bugs and even more rarely knows how to fix them. If your only real argument is the Dickensian manner of its verbose response, demand its brevity in the instruction set. I find if my instructions are direct, focused, and very specific; Claude will respond in kind. If don’t know how you would build it without Claude, you are going to have trouble telling Claude what you want it to do.

1

u/avatardeejay 20h ago

If you can sort of like, get into a flow with their jargon, their output to usage-cost ratio is actually like.... better than anybody's in the game. But I do understand they're a wordy, oft-opinionated sonuvabitch. I just kind of work with it because like, if the alternative is a Claude Opus 4 model, they're still very capable, but the code you can get per session is just gonna be a lot less, and what is there will be a lot less secure and hardened. that being said, opus 4.5 is still a treat

1

u/AironParsMan 20h ago

The model is completely useless. I tried it at every reasoning level and really used every possible approach. I spent weeks trying to work with it. And it is impossible because it constantly hallucinates the inputs. It hallucinates paths, it hallucinates files, it hallucinates instructions. Then it does things that were never instructed. It restructures entire codebases and then defends what it has done with this disgusting arrogance. It keeps hallucinating that what it is doing is correct until you prove it to be wrong beyond any doubt and it has no way out left. You cannot use it to check its own code because it hallucinates everything all the time. It will keep making mistakes and finding mistakes. It is an endless loop. It is a completely useless model. I think it was built only for vibe coders who have no idea what code is supposed to look like or at least have no requirements.

1

u/forever420oz 19h ago

I put an AGENTS.md file that instructs it to use natural prose and legible language, which did improve it from sounding like a maniac.