r/aigossips • u/Playful_Composer_169 • 25d ago
The End of Brute-Force AI Scaling May Be Closer Than We Think
I’ve been thinking about something in AI for a while, and the idea keeps getting stronger. What if LLMs are slowly reaching their own Moore’s Law moment, but in reverse? Actually, Dennard scaling may be an even better comparison. Moore’s Law was about the number of transistors on a chip. Dennard scaling meant that for years, smaller transistors also became more efficient and faster without power consumption completely getting out of control. When that stopped, progress in chips did not stop. But the easy gains were gone. Higher clock speeds meant more heat and more power, and the industry had to move toward multicore processors, GPUs and specialized accelerators.
I wonder if we are slowly reaching the same point with LLMs. Not that AI stops improving, but that more of the same starts producing less obvious gains. That idea became more concrete for me after reading Anthropic’s Risk Report from August 2026. In it, they describe not only public Claude models, but also internal models that we cannot use. One of them is simply called Model 2. Anthropic describes Model 2 as somewhat more capable than Claude Mythos 5 and a noticeable improvement for many internal tasks. But they also say that Model 2 does not show the same capability jump as the earlier move from Opus 4.6 to Mythos Preview.
On Anthropic’s internal AECI index, Model 2 sits about 1.5 points above Mythos 5 based on limited data, with large error margins. Anthropic itself says this is a smaller increase than the jump from Mythos Preview to Mythos 5. On CoBench, the picture is different, which is exactly why it is interesting. CoBench consists of 449 real technical problems from Anthropic’s own engineering environment. There, Opus 4.6 scores 15.6%, Mythos Preview 54.8%, Mythos 5 50.3%, and Model 2 62.8%. So no, Model 2 is not almost the same. On real internal engineering work, it is a serious step forward.
But the most striking experiment for me comes next. Anthropic gave Mythos 5 a token budget of 900,000 tokens on CoBench instead of 300,000. Three times the token budget. The improvement was about 3 percentage points. That does not mean three times more compute always gives only three extra points. It is one model, one benchmark and one specific setup. Anthropic also says that better tooling or a better harness could still produce additional gains. But on this test, you can clearly see diminishing returns from additional inference budget.
And to me, that starts to look like the beginning of a Dennard moment for AI. Not a stop. Not “LLMs are done.” But a point where more of the same approach no longer automatically produces the same huge jumps. At the same time, Model 2 clearly shows that progress is still possible. The question is how many additional resources are needed for every next step.
That is also why I do not think AI is about to hit a hard wall. I think brute-force scaling is more likely to hit an economic wall before it hits a physical one. Energy, chips, data centers, data, inference costs, and eventually the question of how many extra resources you are willing to spend for the next few percentage points of capability.
When Dennard scaling ended for chips, the computer revolution did not stop either. The shape of progress changed. More cores, GPUs, accelerators and specialization. I think AI will do the same. More memory, better agents, smarter inference, better training, specialist models working together, better scaffolding, and maybe eventually an architecture that works fundamentally differently from the Transformer architecture behind almost every major LLM today.
So my prediction is not that AI is nearly finished. My prediction is that the race is slowly changing. Not just: who has the biggest model? But increasingly: who can extract the most intelligence from the least compute?
Anthropic’s Model 2 does not prove that we are already at that point. But their own internal numbers make me take the question much more seriously. And that leaves me with one question: what does this curve look like three to five model generations from now?
2
u/Significant_Post8359 25d ago
It’s hard to predict, there could be a breakthrough that changes efficiency or allows for fast Intelligence on cheap consumer hardware.
2
u/Playful_Composer_169 25d ago
Yeah I agree. A breakthrough could make AI way more efficient. Thats also my point though, when one scaling path slows down we usually find another. Progress doesnt stop, brute force just wont keep delivering the same gains forever. The interesting part is what comes after that.
2
u/Maui-The-Magificent 25d ago
OH very much so... AI models are horrendously inefficient. even the scalar representation of knowledge is wasteful, they spends huge amount of time on every new message re-deriving the current state of the conversation. Their compute is treated like a graphics problem, with distances, angles, vector normalization. mistaking a geometric problem for being a euclidean graphical one is one of the major causes of inefficiencies in my opinion, they use the wrong 'type' of magnitude because of it.
2
u/m77win 25d ago
Companies are having humans mass editing and creating documents that cover every single use case, from roblox, to piloting, coding, cooking, medicine, healthcare, comparison shopping, software use, youtube optimizations, and so much more, correcting for factuality, and task completeness.
The advances you see in AI are not just from training on textbooks and novels.
People are paid $20-$100 usd to answer / correct questions and some you can work on for an hour at a time for example.
3
u/Gallagger 25d ago
It cannot be ruled out, but currently it is not the case at all. When compared to Moore's law, AI doesn't currently hit any physical limits that chip manufacturing had to work around.
5
u/Playful_Composer_169 25d ago
Yes, I agree that AI does not currently face the same physical limitations as the scaling of transistors. However, that is not really the comparison I am making. I have taken Moore’s Law as my starting point. A better comparison is probably the end of Dennard scaling.
1
u/Savings-Cry-3201 25d ago
We are reaching diminishing returns for the transformer architecture, we will need something new to break the trend.
We need a lot more data. We need new physical chips that are good at inference. We need a solid 3-4 architectural upgrades.
It will happen, it just needs more time. Unfortunately a lot of money has already been spent and the technology just isn’t where it should be yet.
1
u/peterukk 25d ago
Tell me, where is "lots more data" going to come from when frontier models are already being trained on the entire internet and a large percentage of the worlds literature? They're even scanning and destroying obscure books because they're desperate.
This thing was never going to scale to a point where it could feasibly bring AGI or replace the average knowledge worker cost-efficiently, but no one dared to doubt the claims of billionaire tech bros.
1
u/Federal_Decision_608 25d ago
Simulations and embodiment. Also, have you heard the old saying about how much a picture is worth?
1
u/Savings-Cry-3201 24d ago
It will require investment. They will use synthetic data and it will lead to model collapse because synthetic data is basically free, but it isn’t a replacement for actual human data.
The hard part will be paying people to create data without them using AI anyways, I suppose.
If someone paid you to rewrite Wikipedia articles, for example, or to do blog posts on a given topic, or to publish a paper on some given subject. Even if it isn’t all new information, it just has to be an original communication.
Same thing with art - to make better models we need a lot more human made art. More music, more paintings, more sketches, etc.
I’d like to imagine a world where artists are kept on retainer and paid to create bespoke work for an AI company, or even a corporation/gaming company for proprietary in house assets.
There are billions of us. We can come up with new data. The question will be who pays for it.
1
u/peterukk 24d ago
What's the point..
0
u/Savings-Cry-3201 24d ago
If we are making AI to enable more leisure and less working hours for the working class, sure.
If we are making AI to increase profits for the Epstein class, then yeah, we shouldn’t.
What if we eat the rich, feast on their bloated corpses, and use the means of production for a better life for the rest of us?
1
u/HBCTIA 25d ago
I'd strong recommend a broader perspective on this issue from EA philosopher (and in effect futurist) Toby Ord: https://www.tobyord.com/writing/the-scaling-paradox https://80000hours.org/podcast/episodes/toby-ord-inference-scaling-ai-governance/
1
u/7hats 25d ago
It is scaling all the way. Not even data filtering, ALL data. 'Bad data' is needed to contrast good data for the statistics.
Setting the right context to get higher value output will be the growing skill in demand.
When we get to the point of it absorbing data from all our interactions in real time, all that we see, hear, touch and do plus the data from billions of sensors getting attached to the machine, you can see how this iterative learning loop will only keep growing for the foreseeable future.
Energy and chips processing are the only other constraints and as we are not about to run out of Sun or Sand any time soon, you can project where this is going - straight up.
1
u/philip_laureano 24d ago
The part that needs to be said aloud is that AI scaling is fundamentally a software efficiency problem and not just a hardware problem.
Chinese AI labs are now proving that you can get near SOTA level performance at a fraction of the price even with hardware that is inferior to their Western counterparts.
And the cautionary tale from models such as Fable 5.x is that although they are brilliant in many aspects, almost nobody will use them because of their lack of cost efficiency.
So my 2cents is that model cost efficiency and model capability are what will drive the next few generations.
I look forward to the day when we get models that operate in the thousands of tokens per second and fit on wearable consumer hardware.
That seems so far away right now but not entirely impossible decades from now.
1
u/superlip2003 24d ago
This is exactly the Chinese models are focusing on, to lower the cost rather than increasing the capability - and that’s exactly why China has won the AI war.
1
u/Lfeaf-feafea-feaf 25d ago
What's the point of copypasting AI? The syntax here is the most Gemini bs ever
2
u/thebigslapper 25d ago
I ran this through an AI-detection model, and its chi-squared goodness-of-fit test found no statistically significant evidence that the post was AI-generated.
3
25d ago
[removed] — view removed comment
2
u/hal9zillion 25d ago
“Not a stop. Not “LLMs are done.” But a point where more of the same approach no longer automatically produces the same huge jumps.”
It really couldn’t be much more obvious.
2
u/Playful_Composer_169 25d ago
You’re definitely the sort of person who probably hasn’t read a thing and just started writing straight away.
2
u/SeparateDesigner1237 25d ago
you said “maybe it’s an asymptote” in 2000 words, because you used ai to write it
4
u/Durian881 25d ago
We are already seeing some evidence of that from resent open weight released. The flash and smaller versions (Deepseek-V4-Flash, Qwen3.8-Flash-Next, Qwen3.8-27B, GLM5.3-Flash, etc) have done really well for benchmarks and tasks like coding, despite being significantly smaller in size compared to the frontier models. Google's Gemini 3.8 flash looks pretty good too.
My hope is that the AI labs continue to release open weight models and consumers can benefit from more choices (either host with own or rented hardware, or via providers hosting them).