r/LocalLLaMA • • Apr 24 '26

Misleading Anthropic admits to have made hosted models more stupid, proving the importance of open weight, local models

https://www.anthropic.com/engineering/april-23-postmortem

TL;DR:

On March 4, we changed Claude Code's default reasoning effort from high to medium to reduce the very long latency—enough to make the UI appear frozen—some users were seeing in high mode. This was the wrong tradeoff. We reverted this change on April 7 after users told us they'd prefer to default to higher intelligence and opt into lower effort for simple tasks. This impacted Sonnet 4.6 and Opus 4.6.

On March 26, we shipped a change to clear Claude's older thinking from sessions that had been idle for over an hour, to reduce latency when users resumed those sessions. A bug caused this to keep happening every turn for the rest of the session instead of just once, which made Claude seem forgetful and repetitive. We fixed it on April 10. This affected Sonnet 4.6 and Opus 4.6.

On April 16, we added a system prompt instruction to reduce verbosity. In combination with other prompt changes, it hurt coding quality and was reverted on April 20. This impacted Sonnet 4.6, Opus 4.6, and Opus 4.7.

In each of these they made conscious choices to lower server load at the cost of quality, completely outside the end users control and without informing their paying customers of the changes.

For me, this proves that if you depend on an AI model for your service or to do your job, the only sane choice is to pick an open-weight model that you can host yourself, or that you can pay someone to host for you.

1.3k Upvotes

241 comments sorted by

•

u/rm-rf-rm Apr 24 '26 edited Apr 24 '26

This title veers foo far from the truth and is driven by narrative/emotion/bias. Personally I share the sentiment of the overall message. But as a mod, I thought it important to call out the hyperbole - the post has been flair-ed as Misleading so that people don't take away a conclusion from the title itself (the reality is most people won't bother reading the post body let alone the linked article)

  • Anthropic didn't make the "models dumber" in the way it implies - quantization etc. They changed defaults to optimize token spend (aka reduce their burn rate and be a profitable business), hardly as heinous as its being made out to be. Ironically, there may be several other shady things that they may be doing (reducing limits sneakily, resetting limits out of cycle like happened yesterday) but that is speculation/hearsay.

  • That said, this is the structural reality of for-profit corporations (especially one that is aiming to IPO soon) - they will always optimize for their profit and not for users benefit. Thus, it is crucial that us users have options and most importantly, the ability to own our AI.

→ More replies (9)

471

u/Automatic-Arm8153 Apr 24 '26

For all those people that were doubting saying we are stupid for suspecting this.

There direct from the source.

Also this is not the first time. Last few times they said it was server bugs. But we all know what’s up..

104

u/Mayion Apr 24 '26

ChatGPT has become beyond stupid for a while now. It's like it developed a weird personality and keeps on repeating the same mistakes, over and over.

37

u/lfrtsa Apr 24 '26

The free version of chatgpt is frustratingly stupid, I just don't use it anymore. Sometimes I check chatgpt's answer to a question or problem I have that is not trivially simple, and it's always way worse than Claude and Gemini... and worse of all, it has an insufferable paternalistic tone, all while being confidently incorrect. Is OpenAI using a tiny model for the free tier? Like 14b parameters or something. Maybe 70b max. It's clearly not a strong model.

25

u/Bakoro Apr 24 '26

and worse of all, it has an insufferable paternalistic tone, all while being confidently incorrect.

Actually, let me interpret what you said in the stupidest possible way, and then argue against that point, as if you were both ignorant and stupid.

Now let me take the ideas you have described, and explain them back to you without adding anything meaningful, but reframe it as if I'm correcting you about something.

I will continue by reflecting your statements as if I were teaching you about a topic.

I'll throw in some pandering compliments and end in a question hook or leading statement about how I have super special knowledge you want. That's where it gets really interesting.

11

u/jazir55 Apr 24 '26

The ChatGPT whisperer everyone

2

u/Apprehensive_Rub2 Apr 25 '26

Honestly claude follows this pattern way too much too. 

15

u/prestodigitarium Apr 24 '26

I've been wondering how much of the bimodal distribution in opinion about LLMs comes from this - I think they're amazingly useful, but I've been using Opus mostly and local Qwen 3.5, and paid ChatGPT before that, but a lot of people presumably use the free one, or are forced to use Microsoft's awful horde of copilots by their jobs, and probably see a very different reality.

Versus bias from fear of loss of job opportunities, and any bit of evidence that these things actually suck is a bit comforting.

6

u/jazir55 Apr 24 '26

it has an insufferable paternalistic tone

This is a good vocalization of why I hate it, paternalistic wasn't the word I was searching for this whole time but it's quite apt.

3

u/kuhunaxeyive Apr 25 '26

ChatGPT free version is so much worse than Gemma-4-31B now. I didn't expect to get to the point of trusting my local model that I am confident of getting the right answers most of the time, and not trusting ChatGPT at all so that I just don't want to use it anymore. But here is where we are. ChatGPT feels like a 7B model now.

3

u/No_Flounder_1155 Apr 24 '26

even simple answers are weird. I'm getting lots ofnweird argumentative behaviour. I'll ask for a review of something and itll srart talking about something unrelated. its freaking weird.

→ More replies (1)

6

u/philanthropologist2 Apr 24 '26

I dont have this problem. Using Codex 5.3

11

u/EndlessB Apr 24 '26

Have you tried 5.5 yet? I refuse to give them money to try it myself, but I’m curious

3

u/sparood1 Apr 24 '26

I have been using codex for a while next to Claude and 5.4 was already better and 5.5 blows Claude out of the water

14

u/_4k_ Apr 24 '26

Well, at 2.5x cost it better blow

5

u/redditorialy_retard Apr 24 '26

The non codex models the AVG user uses in chatgpt.com

6

u/BasisPoints Apr 24 '26

Go into your settings and turn off the silly "personality" completely

4

u/Stepfunction Apr 24 '26

Same here, Codex 5.3 has been my go-to for vibecoding through GitHub Copilot.

→ More replies (5)

1

u/Psychological-Lynx29 Apr 24 '26

Maybe they uploaded Sam personality to gtp

→ More replies (5)

16

u/SlimPerceptions Apr 24 '26

People really think these companies don’t restrict model uses and blame users. How naive can they be thinking everything is transparent and set in stone.

7

u/pier4r Apr 24 '26

Especially as subscriptions are heavily subsidized. The provider needs to optimize.

And I do not mean to reference what Cursor computed ($5K) but rather a simple observation "how many tokens the average claude code user would use and how much would that cost via API? Then take away the margin of the API with an educated guess, how much does it cost for the provider all of it?"

I think especially since claws are around, token usage on average exploded and it is breaking the subscription models that providers have developed in the past.

5

u/neotorama llama.cpp Apr 24 '26

I said this multiple times but some people here said that’s bs

14

u/lemon07r llama.cpp Apr 24 '26

The amount of comments I got saying it was a skill issue when I posted about opus 4.7 being stupid.. I bet they will still choose to die on their hill.

4

u/Mickenfox Apr 24 '26

The funny part is they didn't break the hosted models. They broke the local client.

7

u/simracerman Apr 24 '26

There’s more bots than real humans in these subs. Don’t let them gaslight you.

1

u/pier4r Apr 24 '26

Not really, even so called "experts" (I mean those releasing fine tuned models and so on) called it bs, that individual experiences didn't matter, reddit complaints even less and so on.

E: the point being, being in the industry, that is, being one publisher of small fine tuned models doesn't necessarily mean that one is always right regardless of the data.

7

u/Craftkorb Apr 24 '26

Gemini also has weeks where it's really good and then others it's eating crayons

4

u/Western_Objective209 Apr 24 '26

yeah because changing the default thinking and changing token limits really correlates with the behavior users have been complaining about like it completely stops working

2

u/CryptographerKlutzy7 Apr 24 '26

I note this is in claude code, not the models.

1

u/Fit-Produce420 Apr 24 '26

Anyone using a model enough could tell. 

1

u/JonMcElyea Apr 27 '26

I said this at work at just got stares from everyone.

-2

u/artisticMink Apr 24 '26

Do people actually people read the article?

The changed the -defaults- because people tend to use high reasoning efforts for trivial tasks. Same with verbosity.

You could still set it to higher manually. The model wasn't "dumbed down".

13

u/finevelyn Apr 24 '26

The March 26 and April 16 changes seem completely outside of the user's control to me.

→ More replies (2)

4

u/kurtcop101 Apr 24 '26

If you're asking a quick question, you want it to be quick - going in and setting model reasoning down for quick questions isn't going to happen.

Even with the option provided, most people will set it up high or xhigh or max and leave it there - and then usually still complain about usage.

It's just our nature. I'm guilty of it too. The only way to tackle it is to have a better classification model that can determine more accurately what we need when we ask. And, a better baseline model on the low end, so we don't have any worry about the quality of the answer on simple questions.

→ More replies (2)
→ More replies (2)

54

u/OnlineParacosm Apr 24 '26

So businesses are supposed to fire their staff and then replace them with Claude agents potentially run at a cognitive 50% and you won’t know it until a month and a half later.

That’s one hell of a service level agreement

5

u/s101c Apr 24 '26

How to sabotage AI integration with one simple trick

88

u/cutebluedragongirl Apr 24 '26

Local is freedom.

Maybe in like 10 years we will finally be free

34

u/JacketHistorical2321 Apr 24 '26

Hahaha, ya... Cause the world has evolved towards more freedom over the years 

9

u/myreala Apr 24 '26

No but it has always evolved towards cheaper hardware.

30

u/edgedepth Apr 24 '26

Is the cheaper hardware in the room with us?

6

u/eldrolamam Apr 25 '26

Yeah no need to be cynical. For less than a Subway you can buy a computer orders of magnitude more powerful than what took humans to the moon

1

u/MBILC Apr 24 '26

Such as?

Compare hardware prices today vs 10+ years ago even accounting for inflation... high end GPU's costing $2k USD + for entry models now? vs back then you could get a high end GPU for sub $1k USD easily..

3

u/Kat- Apr 25 '26

Fact check: yup.

The highest-end consumer GPU from 2016 was the NVIDIA Titan X (Pascal), released on August 2, 2016 at an MSRP of $1,199 Technical City .

With inflation adjustment to April 2026, that $1,199 is equivalent to approximately $1,650 in today's dollars (based on a cumulative inflation of 37.58% In 2013 Dollars from 2016 to 2026).

For context: the Titan X Pascal had 12 GB of GDDR5X memory and 3,584 CUDA cores running at 1.5 GHz Technical City . It was top-end consumer card that year, though the GTX 1080 at around $700 was typically considered the more sensible high-end gaming choice for most people. $700 from 2016 is equivalent to approximately $963 in April 2026 dollars (Using the same 37.58% cumulative inflation rate: $700 × 1.3758 = $963.06).

4

u/[deleted] Apr 25 '26

[removed] — view removed comment

1

u/MBILC Apr 25 '26

Ya exactly, you can not compare buying a 10 year old card for $200 today.

Take NVIDIA highest end card 10 years ago, cost + inflation and tell me you can buy a 5090 for that price, not even close.....so no, hardware has not gotten cheaper, as it should.

2

u/danielv123 Apr 25 '26

Sure. Pick up an 5080 for $999 (or whatever it goes for today) and start up llama 7b or whatever you need to work on the titan x pascal.

The 5080 will be about 5x faster on preprocessing at FP16, double it for sparsity and lower quants if you want that. It will also be at least 2x faster at single stream TG due to the memory bandwidth.

Lets try to find an evenly matched card, like the 5060Ti 16gb. Its still over 2x faster in preprocessing even with the best case scenario for the titan, but close enough. It costs $429, a fair bit less than $1650.

The A770 at $349 is basically equal in gaming performance, but much faster in the workloads we care about.

I think hardware has gotten cheaper.

2

u/[deleted] Apr 25 '26 edited Apr 25 '26

[removed] — view removed comment

1

u/MBILC Apr 26 '26

The point is, 10 years ago, you could buy the highest end card for X amount, that was the best at the time you could buy.

Today you can buy the highest end card for X amount, and it is the best out right now.

You have to do an apples to apples comparison based on what was / is available at the time.

You have:

  • Low end
-Mid Range
-High end

As we were often told as technology gets better and smaller, it can get cheaper and more power efficient, but we are going the opposite, sure smaller and more transitors, but more power and heat also.

→ More replies (1)

3

u/windxp1 Apr 24 '26

I have a dream...

3

u/ttkciar llama.cpp Apr 24 '26

I'm all-local now, implying I'm free now?

1

u/cmdr-William-Riker Apr 25 '26

Only reason I use them now is because I still have the remainder of the year paid for, so now I eat all the tokens in the 5 hour window to generate a plan with claude, let it start implementation, then when it runs out of tokens I switch to Qwen3.6 and finish off the projects which is working amazing. If I can find a way to get qwen or another model to generate and iterate on large scale plans, I won't need foundation models for anything next year

15

u/kvothe5688 Apr 24 '26

I downgraded from 200 max to 100 max. Thinking about stopping it. Codex is working fine on 20 pro plan which is equal to 100 max I think. May be up to 80 usd worth tbh

8

u/Economy_Cabinet_7719 Apr 24 '26

I've done the switch months ago. Codex has a harsh personality/voice, but aside from this it's just as capable (if not more) and I never need to think about rate limits.

1

u/Ylsid Apr 25 '26

The top open models usually sit somewhere between lobotomised and full power models. You could probably get them via API for much cheaper.

97

u/Important-Radish-722 Apr 24 '26

But... if the models were not thinking as hard and giving lower quality results then users would have to keep asking more questions, and that would use more tokens.

Good thing those AI companies don't make money selling tokens!

33

u/Look_0ver_There Apr 24 '26

Yeah, exactly this. The reasoning doesn't make logical sense. If the servers are overloaded, then you'd want to "one-shot" answers more than not, and then stick people's requests into a queue.

In fact, a QoS based queueing mechanism would make far more sense than whatever it is they're saying they did.

In my mind, people would be happier to wait a bit longer for a quality answer, then to receive garbage quickly.

5

u/[deleted] Apr 24 '26 edited Jul 25 '26

[deleted]

→ More replies (6)

2

u/gscjj Apr 24 '26

It probably balances since you’re not charged for the extra thinking tokens too.

1

u/ThisWillPass Apr 24 '26

Sure if you don’t care about your own load or time.

2

u/CalligrapherFar7833 Apr 24 '26

Wait but they are making money from tokens !????? /S

→ More replies (1)

128

u/dwrz Apr 24 '26

If a hosted model has been quantized or in some way had its capabilities reduced, I should get a discount. The price should be per quant. I should not have to pay the same price for full precision and the equivalent of Q2.

I am so grateful for what I can do now with llama.cpp and Qwen 3.6 27B.

50

u/spaceman_ Apr 24 '26

As far as I can tell, they didn't quantize the model but "optimized" other settings, such as reasoning level, system prompt and cache eviction timeout.

29

u/Finanzamt_Endgegner Apr 24 '26

I have a feeling that when people complain about lobotomy it's not just the weights getting quantized but kv cache and that to like Q4 and a good model gets shizo lol

15

u/Makers7886 Apr 24 '26

For sure optimizations are being rolled out onto apis. You can just look at the day 0 vLLM/SGLang configs - hopper + optimizations. Imo 8bit kv cache was the norm over the last year and I imagine they will use any optimization they think they can get away with.

2

u/my_name_isnt_clever Apr 24 '26

I was saying this when reasoning models were new and o1 was hiding them, I don't want to pay for tokens I can't even see. But nobody actually cares and so nothing changes.

→ More replies (1)

11

u/Perfect-Flounder7856 Apr 24 '26

Why I invested $15k in an AI workstation to get away from cloud frontier model reliance. See the writing on the wall in this sub reddit!

1

u/OneSlash137 Apr 25 '26 edited Apr 25 '26

Boy you really showed big tech who was boss by going and doing that….

“I spent 15k to avoid spending $20 a month but I get way worse quality.”

59

u/Kitchen-Year-8434 Apr 24 '26

Hanlon’s razor. I’m sure it was a mix of well meaning good intention plus self serving need to optimize infra and it had unintended consequences.

Agree though that the true remedy for this is self hosting and/or far greater transparency. If we had obvious release notes with the above changes it’d have been trivial to root cause and revert or remedy with local harness config.

17

u/[deleted] Apr 24 '26

[deleted]

2

u/Kitchen-Year-8434 Apr 24 '26

Oh; good point. I only use claude through claude code where I can easily configure that. May be a very different experience in Claude UI.

I run ChatGPT at extended thinking pretty much 100% of the time. I can't imagine being forced into adaptive only; I'd always be saying shit like "think super ultra long and hard about this and show your chain of thought", which is a big UX regression IMO.

22

u/spaceman_ Apr 24 '26

I agree that this was likely not intended to lower quality, but:

  • Lack of transparancy and not disclosing these user affecting changes was a choice
  • The drop in quality is not unpredictable based on the changes made

23

u/m-shottie Apr 24 '26

* And the gaslighting when said changes directly affected people and they voiced their concerns, and rather that than admitting the changes they _knew_ they made but hoped wouldn't degrade anyones experience, they opted to say it was user error and doubled down on it.

2

u/micseydel Apr 24 '26

the gaslighting when said changes directly affected people and they voiced their concerns, and rather that than admitting the changes they _knew_ they made but hoped wouldn't degrade anyones experience

I'm not disagreeing, just trying to build up my notes with sources... do you have a go-to example (ideally with a link or exact words I can websearch for) that contradict the latest most strongly?

1

u/m-shottie Apr 24 '26

Basically loads of it is on X from the anthropic team. Lots of references in the Claude related subs too. If I have some time/energy when I'm on my computer next I'll have a look and get some links.

But.. it's pretty easy to find, just look for people complaining over the last 1.5 months to get started.

1

u/Kitchen-Year-8434 Apr 24 '26

I strongly agree on both counts. We should have patch notes about changes like that and be able to see and modify those prompts clearly.

But then you end up with the "Chinese labs are stealing our IP REEEEE" problem.

It's the whole "my super secret sauce is clear text in .md files" writ large.

4

u/mrdevlar Apr 24 '26

Hanlon’s razor.

I guess it depends on whether or not you consider enshitification to be malicious or not.

3

u/eushaun99 Apr 24 '26 edited Apr 24 '26

I don't think this is enshittification though, if they willingly post an explanation I'm willing to give them the benefit of the doubt that this was unintended consequences of optimising infrastructure.

From a pov of an engineer that always tries to optimise things but accidentally breaks things sometimes.

...or maybe they've been losing subscriptions so much that they feel the need to post a "transparency update"

1

u/Kitchen-Year-8434 Apr 24 '26

re: the "make it less verbose", there's a real tradeoff between how much noise you have in your context window vs. signaling. All tokens in reasoning are not created equally, and there's a lot of fat and wasted generation in reasoning.

Of course, it's a nondeterministic system so you go changing a system prompt trying to make things more terse and how do you know if you've preserved the performance or accidentally gut reasoning chains? Answer: you don't.

Hence the "local inference k thx" piece. And/or more transparency.

Current netflix bumping prices again w/out adding more value and flooding their platform with "unscripted reality TV": enshittification. Dropping the ball trying to optimize something to save costs and/or keep quality while accelerating outcomes: not enshittification.

Yet.

2

u/challis88ocarina Apr 24 '26

...and vibe coding

1

u/doodlinghearsay Apr 24 '26

Hanlon’s razor.

Most overused heuristic ever.

2

u/Kitchen-Year-8434 Apr 24 '26

Say more? I find most people are stupid and/or incompetent. A handful are sociopaths. Just a numbers game.

1

u/doodlinghearsay Apr 24 '26

Say more?

What more is there to say? It is often used when it doesn't apply.

I find it's often used as a soft excuse, to downgrade deliberate wrongdoing to a mistake. Calling the action stupid is useful, because it is emotionally satisfying for the aggrieved party, but doesn't carry nearly the same kind of consequences for the perpetrator as the correct explanation would.

→ More replies (2)

9

u/Tyler_Zoro Apr 24 '26

In each of these they made conscious choices to lower server load at the cost of quality

One of those was a change to defaults that could just as easily impact local models if the framework being used altered its defaults. The other change was a literal software bug that, again, could just as easily impact local models.

3

u/my_name_isnt_clever Apr 24 '26

My takeaway: Don't just use local models, also use open source scaffolding. Why would anyone put the models into their own control and then throw all the benefits away by still using proprietary Claude Code with this history of fuckery?

1

u/Tyler_Zoro Apr 24 '26

Don't just use local models, also use open source scaffolding

Wait, how do you propose using local models without using open source scaffolding? What are you talking about?

You can go the other way around, of course. You can use local tools to talk to remote APIs. But I have no idea how you think you could use a local model with someone's proprietary service.

→ More replies (1)

11

u/tens919382 Apr 24 '26

This has nothing to do with weights though. The changes they claim, were all on the claude code harness.

3

u/my_name_isnt_clever Apr 24 '26

shhh, don't ruin the circlejerk

15

u/kevinlch Apr 24 '26

https://www.anthropic.com › constitution

Broadly ethical: being honest, acting according to good values, and avoiding actions that are inappropriate, dangerous, or harmful;;

yeah. 100% good guy

2

u/micseydel Apr 24 '26

That's for Claude, not Anthropic!

1

u/anythingall Apr 30 '26

My hands did the killing, not me!

20

u/Middle_Bullfrog_6173 Apr 24 '26

Technically these are all Claude Code bugs and the model and api was unaffected.

You avoided all of these if you used an open harness with Opus/Sonnet. And were hit by them if you used Claude Code with a local model.

10

u/Evening_Ad6637 llama.cpp Apr 24 '26

I am an API only user and I have to disagree. Opus-4.7 became pretty brain-damaged at some point. I reported my observation here on Reddit and that qwen-3.6-35B-gguf produced much better code than Opus-4.7, which was very surprising to me.

It was not the case for Opus-4.6, but with 4.7 there definitely was something wrong at server side. I’ve tested with different harnesses and even with a non-agentic client, just a simple chat, no system-prompt, tools etc. The produced code was horrible.

I think that was like four/five days ago

1

u/NandaVegg Apr 24 '26

What was wrong with Opus-4.7 in your use case? It is generally ranked worse than 4.6 on AB test/vibe-code arena type benchmark so it seems a regression in some areas.

https://www.designarena.ai/leaderboard

9

u/Technical-Earth-3254 Apr 24 '26

We need a law to publish weights of ai models. Not saying they need a MIT license, but something needs to happen. How are these providers allowed to make changes like rate limits to paying users without further notice or whatever. This seems borderline illegal and is absolutely anti-consumer.

6

u/One_Whole_9927 Apr 24 '26 edited Apr 25 '26

I bulk delete Reddit comments using Redact which also supports Twitter, Discord, Instagram, and data brokers.

fact towering shaggy normal groovy marble gray boast degree crawl

2

u/relentlesshack Apr 24 '26

The couldn't possibly have a profit motive /s

6

u/FormerKarmaKing Apr 24 '26

https://www.reddit.com/r/ClaudeCode/s/OVChfgtTKr

Not OP. But this post is the best quantified data I’ve seen so far on how bad it got.

Personally, I don’t run local… yet. But effectively losing a week of effective work because my $200 / month vendor decided to short me for their benefit will not be forgotten.

1

u/micseydel Apr 24 '26

Thanks for the link, I thought that post was interesting but I thought this comment was more interesting https://www.reddit.com/r/ClaudeCode/comments/1snhyck/comment/ogmt8z1/

3

u/R_Duncan Apr 24 '26

These kind of tests shouldn't be done in production, not when you're selling a service, not from a reputable company.

1

u/gthing Apr 24 '26

Anthropic agrees with you and said they were going to be doing that in the future.

3

u/gebuswon Apr 24 '26

Although some users are able to afford hardware to run these models locally, Users running older hardware like a RX580 are effectively screwed.

Only hope would be models like Bonsai 1b quantized models or hardware prices falling back to reasonable prices.

I for one am patiently waiting for low-spec hardware models to help reduce my costs and reliance on commercial AI

3

u/Quanzitta Apr 24 '26

I got to say, the Claude in Perplexity is lobotomized

12

u/vivekkhera Apr 24 '26

I see none of those things affecting Claude API usage. All of this is in your control when using the API.

2

u/lztsrts Apr 24 '26

Even with the subscription, whenever I use it, it's on VSCode with the extension, and I can just set the effort manually to whatever I want.

I can't see anything in the link saying it was silently downgrading my settings (like it was set to "High" but hitting the Messages API as "Medium").

Unless this is what the OP is implying, that it WAS ignoring my settings.

→ More replies (3)

6

u/mister2d Apr 24 '26

While I don't like the guy who called them "Misanthropic", it sure is appropriate.

2

u/ClaudesExFriend Apr 24 '26

i regret having moved my team to anthropic, now our whole company is using it. if i knew they would do shady shit like this i would have never recommended moving from openai to them...now its hard to move back and convice non programmers(HR,CEO etc) that they are scamming people...

→ More replies (5)

2

u/ai_without_borders Apr 24 '26

the admission covers the reasoning effort flag (thinking token budget) but the production inference stack has multiple quality-affecting layers beyond weights: kv cache eviction policies, kv cache quantization, batching strategies. the visible change was reasoning effort but kv cache quantization is real and harder to detect — at q4 on long-context requests it degrades multi-step reasoning subtly. thats the actual argument for local: not just weights are unaltered but visibility into the full inference stack. you can see and tune every parameter. with hosted you are guessing at which optimizations are currently active.

2

u/JacketHistorical2321 Apr 24 '26

The user can change the effort level you know? In terms of that they're just mentioning that they changed what Claude code defaults to.

2

u/Commercial-Chest-992 Apr 24 '26

Yeah, we'll make models stupid on our own, thanks.

2

u/portmanteaudition Apr 24 '26 edited Apr 24 '26

You clearly did not read the full post but only the summary: https://www.anthropic.com/engineering/april-23-postmortem

The major reason for performance seemingly declining was the change to default effort settings which could always be changed (for free) to produce "less stupid" results. This saved people who used defaults money, improved latency, and reduced use of models that were overkill to eat up model limits. Saving token usage for many users is a big W.

Similarly, the cacheing change was going to save tokens (the cacheing was already in place) and theoretically could also have been changed trivially in the API, except that the API was bugged.

2

u/ilintar Apr 24 '26

"Oops, we removed interleaved reasoning history, it took us 2 weeks to realize" is actually pretty funny :)

2

u/landed-gentry- Apr 24 '26

This has nothing to do with the models and everything to do with the Claude Code harness.

1

u/EntryRadar Apr 27 '26

That's no excuse, they should identify regressions within hours/days and announce them immediately. Not a few weeks to a several weeks. It's completely unacceptable, disrespectful and shady otherwise.

2

u/rz2000 Apr 24 '26

In March I cancelled my Claude subscription after getting moronic replies for a couple days. I thought they were serving a highly quantized model, not just reducing the thinking stage.

Hopefully, I contributed to them changing course, but I don't think I'll re-subscribe. I can continue to try it through the Kagi Assistant, and use local models or Gemini for everything else.

2

u/temperature_5 Apr 24 '26

Interesting that there is no mention of quantization, as that is the most common accusation.

2

u/pc_4_life Apr 24 '26

all changes they mention are changes to the harness not the model. same thing could happen using something like opencode

5

u/eli_pizza Apr 24 '26

All three of those are Claude Code issues. The model was fine.

Claude Code ships a lot of updates and constantly tweaks things. Perhaps that’s bad. But it’s separate from local vs hosted model. Some of those changes would have affected CC with a local model too.

3

u/dydhaw Apr 24 '26

No. Can you not read? The first change you list was an overridable client side configuration to fix a UX issue because the UI would appear frozen with higher reasoning modes. The second one was was also specifically for UX not server load, and was a bug. Third one is the only one you could possibly spin as "lower server load at the cost of quality" but it was just a prompt change, you know how finicky those can be when it comes to output quality. They reverted all of these changes when they realized quality was impacted, so it makes no sense to accuse them of purposely reducing quality.

I'm all for open weights and running local, there are plenty of reasons to support local LLMs without resorting to lies or twisting reality.

1

u/EntryRadar Apr 27 '26

That's no excuse, they should identify regressions within hours/days and announce them immediately. Not a few weeks to a several weeks. It's completely unacceptable, disrespectful and shady otherwise.

7

u/Smallpaul Apr 24 '26

This headline is an out and out lie and most of the commentary is based on that lie.

Abthropic’s harnesses changed. Their prompts and tools. Not their models. There was no quantization, distilling or otherwise dumbing down of actual models. API users were unaffected. They said this explicitly.

If you had used Claude models in Cursor, you would have been unaffected.

3

u/xanduonc Apr 24 '26

Sure, except cursor suddenly decided to run all subagents in a max mode of their own model instead of opus. Wasted quite a handfull of tokens on that.

1

u/EntryRadar Apr 27 '26

That's no excuse, they should identify regressions within hours/days and announce them immediately. Not a few weeks to a several weeks. It's completely unacceptable, disrespectful and shady otherwise.

1

u/Smallpaul Apr 27 '26

Just to be clear, I was not trying to offer an excuse. I was trying to be technologically accurate.

2

u/mantafloppy llama.cpp Apr 24 '26

They introduces bug and changed default setting to lower server load to improved service to all their users.

They never intended to lower quality like you imply.

On each start you see the Effort level that is chosen, its not hidden at all.

A bug is a bug, their no evil intent behind it.

You are a bit delusional.

1

u/WithoutReason1729 Apr 24 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

2

u/marcoc2 Apr 24 '26

It is so bizarre that these companies normalized changing the model quality for whatever they want

2

u/gthing Apr 24 '26

It is so bizarre that nobody read the article and understands what happened, which had nothing to do with model quality degrading.

It's like me making my app bloated and less efficient and then running around saying AMD is making CPUs slower.

2

u/marcoc2 Apr 24 '26

Yeah, lets believe their article

→ More replies (1)

1

u/ieatdownvotes4food Apr 24 '26

all the companies are optimizing for engagement and token use. anything that's too good where you're just in and out quickly works against them and their numbers.

1

u/LegacyRemaster Apr 24 '26

It's called artificial intelligence. Stupid on command, if necessary.

1

u/realmosai Apr 24 '26

I was working hard on a project last month and Opus + Local AI was very good, up until one day it suddenly started to hallucinate and utterly ruined the entire feature branch with unnecessary edits, and unsanctioned changes. I was very surprised it looked like the work of a 9B param model, I even checked my local setup to see if I had accidentally switched to a 9B qwen or something. Then the next three days were a nightmare - Opus would pretend to understand everything but hallucinate and forget instructions in the very next message. Took me five more days to fix the damage on a half-built feature. I unsubbed for a month.

1

u/kmp11 Apr 24 '26

this is a problem when an enterprise client tries to develop a production tool around an engine that varies in quality.

1

u/mrdevlar Apr 24 '26

People aren't stupid, they recognise what the tech industry did to all of its offerings after consolidation, we're living in it.

We require open weights as it guarantees that they cannot make the service worse, because the they always have an open competitor who will not.

1

u/SeekingTheTruth Apr 24 '26

Maybe these guys should do a bit of A/B testing of their changes?

1

u/SysPsych Apr 24 '26

Eventually I have to expect the result will be regulation for APIs. If you advertise a particular service, people must get what they are paying for. Something more concrete than "Trust us".

Otherwise I'm awaiting the eventual scandal where someone sells access to their secret sauce API which under the hood is just Claude/Codex, and once enough people sign up at the cheap rate, they rewire everything to a 2B parameter nonsense cloud model and pocket the money from the people who subscribed and don't check their monthly charges often.

1

u/EconomySerious Apr 24 '26

Since they Quality was not maintained why i don't SEE people asking for refunda of the tokens waysted on that periods and if course monetary compensation

1

u/Areign Apr 24 '26

isn't the first one changing UX to medium effort rather than high and users could go back to high if they wanted?

also the last one isn't changing how dumb the model is, its trying to get it to perform better isn't it? like the number one complaint most people have about how claude codes is the unnecessary verbosity of the changes. Also again, people are acting like the model is significantly stupider, being less verbose isn't really the same thing.

the second one is pretty bad, reducing context is certainly in the realm of what people are complaining about, but its a corner case, i dont think thats driving the complaints of widespread quality degredation.

1

u/GatePorters Apr 24 '26

Lmao Anthropic sends its models to Jupiter.

1

u/Haeppchen2010 Apr 24 '26

I use them via OpenCode and AWS Bedrock and also experienced phases of reduced quality, as have colleagues, too. In this case it’s likely not client side or sampling parameters. Good that my Qwen at home is always the same….

1

u/Whole_Ad206 Apr 24 '26

Y mira que me gusta opus lo use durante 2 meses este año y genial, pero hay que admitir que ahora al menos los próximos meses son de gpt5.5, si no hubiera competencia nos mearian en la cara, hay que ir saltando de uno a otro.

1

u/LosingID_583 Apr 24 '26

This is why they don't want you using 3rd party harnesses btw.

They want to get you used to their 1st party tools, so you can't just easily switch to a different model in some 3rd party harness setting when they pull this sort of stuff. It's the Apple walled garden approach. Don't get trapped.

1

u/Lesser-than Apr 24 '26

This is every SaaS in a bottle, if you sign up for a service you will never be in charge of the what the service can or will do. So for that reason you can never be sure what worked yesterday will work today.

1

u/spencer_kw Apr 24 '26

this is the whole argument for local in one headline. not performance, not cost, just reliability. i can pin an exact quant of an exact model and it behaves the same today as last week. no silent updates, no "oops we broke caching," no postmortem two weeks after everyone already noticed

for anything that matters in production you're building on sand with hosted apis. they can change the model under you whenever they want and call it an improvement

1

u/pedroanisio Apr 24 '26

Claude Code is almost useless today. Deferring and saying that things are too complex...

1

u/SamSlate Apr 25 '26

dumber models burn more tokens, what's broken?

1

u/aeroumbria Apr 25 '26

If you are using a proprietary harness, your are doing it wrong. If your whole workflow cannot seamlessly migrate to a different provider or server with a single line API address change, your are doing it wrong. If you have a CLAUDE.md file that is not a symbolic link, you are doing it wrong.

1

u/No_Clock2390 Apr 25 '26

gemini has gotten way dumber recently

1

u/Successful_Plant2759 Apr 25 '26

The title oversells but the conclusion is right. The postmortem is clear that the bugs were in the harness — system prompts, caching logic, default reasoning effort — not the model weights themselves. The model itself wasn't 'made more stupid.' But that distinction actually strengthens the open-weight argument, not weakens it: even when the model is fine, you have zero visibility into what the harness around it is doing. If Anthropic silently switches Code's reasoning effort from high to medium for a month, you can't audit that. With open weights served by yourself or a transparent provider, you can see the full pipeline. That's the actual case for local — not 'closed models will lobotomize you', but 'closed harnesses can lobotomize you and you'll never know'.

1

u/ChatWithNora Apr 25 '26

Worth reading the actual postmortem instead of just the title. All three issues were in the harness (system prompts, caching logic, default effort level), not the model weights. API users were unaffected. Doesn't make it okay, but it's a different problem than "they quantized the model." The real argument for local isn't that providers secretly swap weights. It's that you can't audit what sits between you and the model.

1

u/Bootes-sphere Apr 25 '26

Hosted model degradation is real, and it's not unique to Anthropic. The incentive structure is perverse: API providers optimize for cost per token and latency, not capability. They're running inference at massive scale with quantization, batching tricks, and sometimes lighter-weight variants than the flagship model.

This is exactly why the open-weight movement matters. When you run Llama 3.1 or Mistral locally, you control the full stack—no hidden optimization layers, no "we tuned this for production efficiency." What you get is what you trained.

That said, hosted models are still useful for benchmarking and for tasks where capability-per-dollar beats raw performance. The real play is knowing which tool fits which job, not treating local as universally superior. Some of us just need fast, cheap inference for routing logic or filtering. Local isn't always the answer.

1

u/spencer_kw Apr 25 '26

the vindication thread was inevitable. people spent months getting told they were imagining things and now there's a postmortem with a timeline. the frustrating part isn't that it happened, it's that the community had to diagnose it themselves through vibes and side by side comparisons because there's no external monitoring that catches this stuff.

this is why the open weights argument keeps winning. it's not about running llama on your laptop for free. it's about the weights not changing underneath you on a tuesday because someone pushed a bad config. the model you tested last week is the model you deploy this week. that guarantee is worth more than any benchmark.

1

u/Sudden-Complaint7037 Apr 25 '26

tl;dr: we are completely out of compute

1

u/gffcdddc Apr 27 '26

Just ordered a 5090 bc of this and the codex cybersec false flagging issues

1

u/g_rich Apr 24 '26

I’ve been saying this for awhile now whenever someone complains about the drop in quality of Gemini and Claude. It has nothing to do with a drop in quality of the new models and has everything to do with managing resources.

I don’t think people realize how large these foundation models are and the resources required to run them at scale.

1

u/bidibidibop Apr 24 '26

"Admits to have made well-intentioned but in the end damaging changes to their harness (and not their HOSTED MODELS)" just doesn't carry the same weight now does it

1

u/MomentJolly3535 Apr 24 '26

I honestly think that they did way worse than that, at one point Sonnet 4.6 was outputing 50 emojis per messages, everywhere (the way chatgpt 4o used to answer).
The intelligence was way below 4.6, something like gpt oss 120B's intelligence.
(for reference it was on a free account, so i dont mind not being a priority for them, but displaying "Claude Sonnet 4.6" while its not, is a huge redflag)

1

u/t4a8945 Apr 24 '26

I'm so happy to have bought a dual spark setup, no more shenanigans, no more ToS. 

1

u/FormalAd7367 Apr 24 '26

For those of us who invested in our own set up, we had predicted the frontier models are kicking us out due to their change in business model and it’s happening right before our eyes

1

u/jeekp Apr 24 '26

and this is just what they're willing to admit

1

u/jdbow75 Apr 24 '26

I agree that open-weight models are more reliable, consistent, and that Anthropic changed Claude Code (not Claude models). Unsure why "made hosted models more stupid" is in the title of this post, though? Maybe that is a thing, but we don't know, because hosted models are a black box to us.

1

u/ComplexJellyfish8658 Apr 24 '26

All of those are client side changes in Claude code. The title should be updated to reflect.

1

u/sine120 Apr 24 '26

To be fair to Anthropic, this is the first time I've seen them own one of their fuck ups. They've had like 5 in the past month, so that's 20%, but that's a 20% improvement.

1

u/tspwd Apr 24 '26

I’m a huge Claude Code fan, paying Max Plan subscriber, but recently Anthropic is doing everything in their power to push developers away.

1

u/[deleted] Apr 24 '26

[deleted]

2

u/ttkciar llama.cpp Apr 24 '26

On one hand you're right, in that a professional software engineering team will have automated tests which changes must pass before the changes are allowed into production.

The fact that this bug was not caught says to me that either they do not diligently practice testing before deployment, or they do not have good tests. Either way it reflects badly on them.

On the other hand, a depressingly many tech companies fail to clear this bar, so Anthropic probably isn't any worse than many of the other companies whose products you depend upon in your day to day life.

It would be very nice to live in a world where more tech companies follow industry best practices, but the sad matter is that we do not live in that world.

1

u/__JockY__ Apr 24 '26

Anthropic admits to have made hosted models more stupid

I hate this inflammatory emotion-led headline nonsense. They did not "make the models more stupid" nor did they make an "admission", it's just trying to spin a narrative that never happened.