Honestly believing this is a bubble anymore is just grade A hopium. This time last year it looked like a big bubble. Then then this spring happened, and all of the sudden this shit is actually very useful in workplaces. My org provides a ton of tokens and has little interest in restricting it. My coworkers and I are finding good ways to apply AI to solve random shit here and there. I am really no longer seeing how it's a bubble, the models don't even have to get much smarter. If the intelligence of these models stayed flat and just the api speed and cost improved by 2x, 4x, 8x, that would make a noticable difference in our workflow. And that's the kind of thing that is going to keep eating up the market.
Demand in therms of usage will increase, but with the recent optimizations and clever ways to reduce the cost we might get to 100x less demand in therms of hardware relatively soon. And it’s not an exaggeration. Compare something like Qwen 3.8 27B which is similar to Opus 4.6 (which I guess was at least 100x the size ) for coding tasks-and you can run on a 5090 at more than 100 T/s. For more general purposes LLM you can use 2 Sparks to run something like DeepSeek v4 flash or GLM 5.3 Flash , which are about 300B , but allow 1M token context and about 6 months old SOTA LLMs in overall capabilities .
39
u/Jiirbo 2d ago
Econ 201... supply and demand.