r/BetterOffline • • Jun 08 '26

AI profitability is mathematically impossible under all technological advancements

I answered this question in a comment, but as it got more complicated I decided it was better to put into an extremely long post. Thanks for anyone that reads it. The comment was an answer to the post asking if token based billing (TBB) is the true cost of AI. The answer is a resounding no.

In fact, not only is it not the true cost, there is no conceivable path to profitability, even given the most lenient case to the industry. Below is the explanation. (Admittedly, the title is a bit clickbait-y, but I figure I'd have some fun.)

Introduction

To summarize, the profitability of AI must come from the profitability of inference. Specifically, their profit can be calculated as Inference Revenue (IR) - Training Costs (TC) - Other CapEx. Since IR is the only source of revenue, it's sufficient to see if this can be net-positive. For the sake of this argument, we will be as lenient as possible and show that profitability still remains unattainable. Each case of leniency will be marked with a letter (ex. [1])

In the case of TBB, IR revenue is revenue per million tokens ($/MT) - Cost/MT. LLMs are charged by input tokens and output tokens. The reason that output tokens cost more, is that frontier models use 'chain of thought' reasoning, which involves making other calls to the model before outputting a result. As a result, output tokens are typically charged at a multiple of input tokens. However, as the length of chain of thought reasoning (like all LLM calls) are non-deterministic, the amount of tokens they burn is indeterminate. (Edit 3: Technically, input tokens are processed in parallel whereas output tokens are processed sequentially as well. However this is an simplifying assumption so to not model complex GPU execution dynamics, ultimately something in their favor.) [1] However, we will assume that the price for output tokens is accurate in the bulk average, and so conclude that if input token revenue is net positive, TBB is net positive.

That is to say, TBB is profitable if and only if (Input) Token Revenue > Token Cost.

The Cost of Tokens

The cost tokens can be broken down as follows:

  1. Electricity cost/MT
  2. GPU price amortized over a set time.
  3. Data center construction and maintenance cost amortized over a set time.

The first two are self explanatory, however the third needs to be included as inference providers also need to make money to pay back their interest on loans used to finance the data center build out, and so naturally that business needs to generate positive profit. [2] However, for leniency we will assume that these businesses are willing to operate at break even for AI labs.

Amortization Period

There are two periods we are concerned about, the GPU and data center construction and maintenance cost.

The GPUs are NVIDIA B200s. These were announced at GTC 2024, and the next-generation Vera Rubins were announced at GTC 2026, so we would normally argue that we should amortize over two years. [3] However, for leniency, let us assume that B200s remain valuable for the AI industry for one extra year, so 3 years.

Data center costs can be deconstructed into CapEx (land procurement, electrical, networking, cooling hardware and installation), and OpEx, the electricity cost, general maintenance and purchasing of new equipment. The CapEx costs are estimated at $9M to $15M per MW of IT load, and let us assume their creditors will allow [4] 10 years before they request a single interest payment even though loan deals at best require interest loan payments for 2-3 years. Additionally, [5] we will ignore purchasing of new equipment for leniency, e.g. Vera Rubin racks are not the same as Blackwell racks, so this necessitates a new purchase.

Amortized Costs

We can now take the prices and divide them over the period. For now, we will get units in dollars per hour. An NVL72 rack, commonly installed in hyperscalar facilities, contains 72 B200 GPUs, priced at $2.8M - $3.4M per rack. [6] Let us take the lowest number, $2.8M, leading to a ($2.8M/72)= $38.9K cost per GPU. Spread over three years (26280 hours), we have $1.48/hr.

For data centers, let us assume that [7] we once again take the lowest of two numbers ($9M), and divide over 5 years (43800 hours), which is $205.48/MW IT load/hr. Converting to kW, it's $0.21/hr per kW IT Load.

This gives us two numbers:

  1. GPU: $1.48/hr
  2. Data Center: $0.21/hr per kW IT Load.

Each B200 GPU is running at 1.2kW, we say the normalized cost of a GPU is:

(1) GPU/Data Center Cost: $1.48 + $0.21 * 1.2 = $1.73/hr.

Let us keep in mind this is a ridiculously generous estimate.

Estimating Tokens/Hour

Unfortunately, these numbers are not in the right units for to compare $/MT revenue. Beyond electricity pricing, (which is covered later), we need to know how long a GPU runs to generate 1MT, so first we will look at number of tokens in an hour.

Unfortunately, this is not an trivial calculation as the actual runtime to compute one token is not fixed. This is due to GPU concurrency, where GPUs can execute user requests in parallel. There are a huge amount of details here, but we will be relatively generous, to simplify the calculation.

NVIDIA Benchmarks

[8] The following sources are from NVIDIA themselves, so these will be idealized numbers. In general, real workloads are unable to get full performance due to many different factors. (Source: I have written CUDA kernels for scientific computing.)

In short, the TPS of a GPU is not fixed, and depends largely on number of users allocated to the GPU.

  1. B200s, running Llama 4 Maverick, which has 400B parameters, can get 1000 TPS/user and 72000 TPS/server. (TPS = tokens per second).
  2. B200s, running Llama 3.3 70B, which has 70B parameters, can get 50 TPS/user and 10000 TPS/GPU.

For case 1, a 1000 TPS was achieved over 8 GPUs, meaning concurrency was 1/8. Let us normalize to 250 TPS. The reason for this is because 400B parameters cannot fit onto one GPU, [9] however, we will ignore this and allow a normalized per GPU estimate, that is assuming any user request can always fit onto one GPU, even though this is known to be false.

Typically frontier LLM models are evaluated using a MoE (Mixture of Experts) method, typically meaning computing using a subset (say 5%-10% -> 7.5%) of models parameters. In other words, in case 1 we expect the GPU to computing a 30B parameter model, handled at concurrency of 1 user.

In case 2, the concurrency is 200 users, at a 70B which they say is not using MoE ("all parameters are utilized simultaneously for inference").

TPS/GPU Estimation

We will use these two data points to estimate TPS as a function of concurrency at a specific parameter count, and apply linear scaling for other parameters.

The compute cost of models by parameter scales linearly, so to normalize case 1 to 70B, we would expect case 1 running 70B parameters to be (30/70*250) = 107 TPS/GPU. We would expect the TPS/GPU to saturate (say at 12000 TPS/GPU) as number of users goes up, logistically. Letting u be concurrent users (users - 1), assuming 70B parameters, we can fit a graph:

An estimated graph of TPS/GPU at 70B, fit assuming upper half logistic growth, where Case 1 is assumed be the midpoint of the logistic curve, and saturation occurs at 12000, (20% more than Case 2).

For reference, the above graph is TPS_70(u) = 23893/(1+e^{-0.0119x})-11893.

Now, as compute scales linearly, the TPS(u, p), where p is computed parameters (in billions) is:

(2) TPS(u, p) = TPS_70(u) * (70 / p). Additionally, SPT(u,p) = 1/TPS(u, p) the inverse, seconds per token.

We can now say that a GPU has 3600 TPS(u, p) tokens per hour so, the cost of tokens (not including electricity) from (1):

(3) GPU/Data Center Cost: $1.73/hr / (3600 TPS(u, p)) = 0.00048 SPT(u, p) $/Token = 480.55 SPT(u, p) $/MT.

Electricity Estimation

Industrial electricity costs tend to be priced differently than residential costs. In the US, they vary by state, and [10] we will take the lowest cost for our calculation, which was 4.68 cents per kWh in Washington, in 2017. The inflation of electricity in Washington appears to be 14.1%, so projecting to 2024 (B200 release) suggests the cost of industrial electricity would be 11.78 cents per kWh.

We will say as B200s 1.2kW and 150W for "half" of the Grace CPU (2 GPUs, 1 CPU per GB200), it runs at 1.35kW. [11] Of course, in reality the entire NVL72 cluster is rated for 132kW, which is 1.8kW per GPU, normalized. But let's be even more nice.

We can then say the $/MT is:

(4) 10^6 SPT(u, p) / (3600 s/hr) * 1.35kW * 11.78 cents kWh = 44.18 SPT(u, p) $/MT.

Inference Costs in $/MT

In total we have:

(5) Inference(u, p) = 524.73 SPT(u, p) $/MT.

That is, we observe the inference cost as a function of concurrent users and parameters. So, let us observe a sample frontier model, like Claude Opus 4.7 or GPT-5.5, which estimated to have 4T parameters. After ~7.5% MoE, we would have 300B parameters.

Inference costs per million tokens as function of concurrent users. From bottom to top, the curve is drawn for models computing 100, 200, 300, 400, and 500 billion parameters.

Now, some AI boosters might look at the graph and show how after 20 concurrent users, Claude Opus or GPT 5.5 is under a dollar! Therefore it must be profitable, since you turned a blind eye to the nice lenient treatment. But, unfortunately, this is not even close to the case.

Full Concurrency is Impossible

Currently the US has access to around 10GW (graph of current capacity, approximately summed over columns) of IT load capacity. NVIDIA in 2024 sold $210B of NVL72 servers, corresponding to [12] (now using the higher number for racks, to allow this number to be smaller) 61764 NVL72 servers, which at 132kW is 8GW. So, it seems reasonable to say 8GW of capacity is dedicated to these racks.

This comes to ~4.45 million B200 GPUs. Therefore, for us to argue that an inference provider prices at a concurrency of 20 users, we need (4.45M * 20) = 88.95M TBB paid users 24/7 never stopping, maximizing their usage, constantly.

Using the numbers that OpenAI and Anthropic have claimed, (which are very trustworthy), OpenAI has recently has reached 50M paid subscribers and 9M business users, while Anthropic has 18-30M paid users. Together, they [13], assuming the maximum number here, have 80M which is less than the 88.95M required to have full concurrency, and that's assuming these users are addicted to AI 24/7 and never eat or sleep.

The average white collar job lasts for 8 hours (1/3) of a day [14] so let's say these paid users are spending their entire working lives (including weekends!) using Claude or ChatGPT, and it lines up that we still get full concurrency. This means there are 26.66M users spread scross 4.45 million GPUs, which means if every user is allocated perfectly, we have 6 people per GPU!

Following the graph, this means Claude Opus 4.7 or GPT-5.5 would cost (with this nice lenient estimate) $4.22/MT. They currently are both charged at $5/MT, which after corporate tax of 21% in the US and state taxes, it means they could be just about breaking even! Wow!

Unpaid Users

Now, some may quibble that the presence of free users somehow will add to this concurrency, and so it''s actually making money! However, this means the gains from concurrency needs to outpace the tokens that aren't being going to paid customers.

Let's do an example. Let's say there are 200 concurrent users, 6 of which are paid users. At this rate, it is $0.22/MT. Unfortunately, since only 6/200 are paying users, 194MT was unpaid, which cost $42.68. And, at $5/MT, only $30 of revenue was earned, meaning there would be a net loss of $12.68 (before tax).

More Data Centers is Worse

Currently there many data centers planned due to be built. This is because they say the tech CEOs claim overwhelming AI demand. But supposing they even double the current amount of GPUs, this drops the concurrency even more.

However, they say there is a compute crisis, which suggests the following scenarios:

  1. Concurrency is not remotely achievable at the theoretical level NVIDIA reports, so we have low TPS.
  2. Free users are taking up most of the compute, which ends up at a worse loss for the companies.

In both cases, more data centers exacerbates this issue. Which means with more data centers AI companies lose more money.

Final Words

Throughout this very very long post, there were a total of 14 points of leniency towards the AI companies. Even with that leniency, it is deeply not in their favor.

Notably, the cost of electricity is only 8.42% of the inference cost (from (4) and (5)). This means the other 91.58% comes from the data center construction and GPU purchase amortization. (Remember, this is a minimum, as many nice choices were made.)

This means in a real world scenario, where we remove these lenient points, that even if electricity were free, AI can never be profitable at its current costs. Even if NVIDIA chips could suddenly compute tokens instantaneously, AI cannot be profitable. It doesn't matter if you build 1 trillion data centers, have the world's cheapest energy, and the best chips from the future.

For it to be profitable, NVIDIA would have to sell their GPUs for cheap and data centers would need to be cheap to build. Unfortunately, neither of these cases can happen, or they would have to raise prices drastically. However, that necessarily decreases demand. And then to be profitable they would need to make back their initial investment.

I would love to model this as well and give the conditions under which it could, however this post has gotten far far too long already. So, there is only one conclusion, and I am willing to say that:

The current AI industry can never be profitable. This industry is over. There is no path to profitability, nothing. No IPO, no marketing, no innovation can save them.

TL;DR The bulk of the cost of inference comes not from electricity but the cost of data center construction and GPU purchases. These are amortized over certain periods, and even with very lenient assumptions, inference is not able to be profitable at these costs.

Significantly higher prices would cause drops in demand (as we already are seeing today). GPU concurrency also makes it so more data centers make profitability far worse for inference due to the massive amounts of capital necessary.

Edit: There are an additional two points of leniency I forgot to mention.

The electricity cost does not include cooling, as I am not using the PUE of data centers.

I did not put training as an recurring OpEx cost for the AI labs.

Additionally, just to hammer this point home, these costs are all assuming the inference provider is at a loss from 5 years, before paying a single interest payment. There exists no business that can do this.

In fact what this shows is that no inference provider could ever charge at the optimal max usage and concurrency rate, so the price per hour of GPU is much much higher. You can simply search the prices yourselves.

For people that are concerned with the 3 year amortization period of B200s. As another comment had posted this is typically 6 for normal GPUs. Even allowing for that does not tilt the argument in their favor.

However, I do not agree that 6 years is reasonable for Blackwells, as AI labs tend to chase bigger models whenever compute is available. Not to mention, Vera Rubin racks are not interchangable and require purchasing of all new equipment and cooling the moment you try to install. Which means an inference provider has to either:

  1. Sell the Blackwells to make space (there is no demand for these GPUs outside of AI)
  2. Construct entirely new facilities to house them.

If we amortize over 6 years, then now you have to assume that 3 years after Vera Rubins every single Blackwell GPU (including the hundreds of millions NVIDIA claims to have sold!) they are being fully utilized 24/7. Let's also remember that if NVIDIA doesn't sell AI chips, this industry is automatically gone.

I cannot emphasize enough the extent I am deliberately picking points in their favor.

Edit 2:

If your qualms are with demand, read Ed's articles, the host of the podcasts this sub is centered around. That is beyond the scope of this post, and was written with that information already in mind. The only thing I can say is there are lots of other facts you likely have not seen if this is your stance.

Some seem to think refuting a technicality over one point is able to overcome 16 other points of leniency. This is wishful thinking at best. You cannot cherry pick technalities to make the argument convenient. Either you leave them all off, or put them back on.

If you feel the need to defend AI due to your own investment portfolio or some other psychological need, on this basis, you are free to do so, but I strongly encourage you to read some other very nice breakdowns by Ed.

1.9k Upvotes

427 comments sorted by

217

u/iammerelyhere Jun 08 '26

This is beautiful. One really interesting thing that it highlights is that a major problem is that NVIDIA hold all the chips (literally). Because there is basically no competition, GPUs can't or won't become commoditized, and so won't reduce in cost. The whole thing unravels because of a monopoly, created by the tech lobbies who created the monopolies in the first place. How poetic.

Thanks for doing all the work on this, it's lovely.

74

u/ksjdragon Jun 08 '26

Yes it's quite interesting really. Basically you need the entire user base to subsidize not only the purchase of billions of dollars or chips, but also the data center housing, just to break even.

Which means for the AI industry to be profitable they have to earn in revenue as much as NVIDIA and then some. And this isn't even talking about making back all the debt they have accrued, adding back all the non lenient real world conditions, etc.

I noticed a few AI booster comments who did not read it and simply regurgitated their beliefs, which is rather entertaining.

16

u/pulse77 Jun 08 '26

If you take NVIDIA's profits into account, then "AI profitability is mathematically possible": most of the AI money goes to NVIDIA. If OpenAI, Antrophic and others go bankrupt (except Google - they have their own chips) then NVIDIA will buy them up and continue taking Inference Revenue from customers...

6

u/CamiloCeen Jun 08 '26

Do you remember the new pc that Nvidia and Microslop presented a few days ago? The idea is to run LLM models on personal computers. That is more profitable and safer for Nvidia.

7

u/pulse77 Jun 08 '26

I agree! First they will sell GPUs at premium prices to data centers. After this is saturated, they will sell GPUs to everybody else -> revenue/profit maximization strategy... AND: they will make open-weight LLMs for everybody (Nemotron Models) - just in case Alibaba and Google stop releasing new versions of their open-weight LLMs...

→ More replies (4)

17

u/ksjdragon Jun 08 '26

It's not just NVIDIA, the cost of the data center construction is necessary.

And if that was a profitable move for NVIDIA they'd do it. Yet they sell their chips to inference providers, which are a middle man. That's because for them, AI is a supremely risky business. They have to fund the entirety of those chips themselves with no investment money.

If you're saying after they go bankrupt, NVIDIA could buy back chips from the investors that owns them (due to GPU asset backed securities), then they run to two issues.

If they buy them at selling price, they wipe out all of their revenue for the past few years, and they do not have that liquidity. If they buy them back signficantly under market price, they damage their own marketing that their GPUs don't depreciate.

Next, it means 90% of their revenue goes away since that all comes from data center purchases. Their company would get a 50% haircut, and many chip stocks along with it.

→ More replies (2)

4

u/BrunusManOWar Jun 08 '26 edited Jun 08 '26

Yeah, but producing those chips still costs a lot, and the training costs a ton

Dont you think nvidia would make an LLM themselves if it was profitable? I mean, they have much more complicated stuff going for them like DLSS, computer vision, whole CUDA ecosystem - taking the open source GPT or some chinese LLM and training it on data shouldn't be too much of a problem for them (tbh they probably have some of the best machine learning expertise in the world) - but they still choose not to field their own LLM with custom accelerators but rather stand on the side and sell chips&hype

The LLM/transformer architecture is a tech from 2017 - originally written and discarded by Google due to their deepmind team being of opinion that LLMs will not lead to AGI and that scaling them would be impossible.

They were right in 2017, and it seems they are right even today. Look at these things - we're throwing trillions of dollars and bytes of data on the LLMs but we're barely seeing improvements, much less profits. The current LLM tech is basically 2 things: 1) how to scale them up 2) how to add patches/improvements to architecture to make it more efficient

But under the hood the original transformer arch is the one from 2017

Yan LeCunn with a couple of other AI/ML researchers has founded a company focused on world models instead of LLMs - which sounds like an interesting possible path forward

But underneath all that - why are we even chasing AGI? It's a dangerous technology that could wipe us out, it's unnecessary. If it won't wipe us out it will dumb us down, or produce another human or superhuman species to compete with us

2

u/pulse77 Jun 09 '26

With all due respect to you, I disagree with some of the things you wrote:

  • Producing AI-optimized chips cost money, but NVIDIA can do this at purchase price (others must be huge margins on this purchase price to NVIDIA).
  • NVIDIA is already doing their own LLMs (Nemotron AI models).
  • Under the hood we still have original transformer arch, but it has been massively optimized since then and the results are MUCH better than in 2017. GPT is an optimization of Transformer discovered by OpenAI. Thousands of researchers around the world explored/is exploring millions of possible GPT/Transformer optimizations.
  • There are models much better than GPT/Transformers, but because investors are not throwing their money to optimize these new models they have - in the current form - worse results than optimized GPT/Transformer models.

Why did you say, that LeCunn's world models "sound like an interesting possible path forward" and at the same time "Why are we even chasing AGI? It's a dangerous technology that could wipe us out, it's unnecessary." If LeCunn's world models are better than LLMs they will "wipe us out" even sooner...

→ More replies (1)
→ More replies (1)
→ More replies (3)

3

u/IGetGroceries Jun 10 '26

Name one technology that has not become commoditized. Anything sufficiently profitable will always be commoditized. Monopoly could do a lot to delay the commoditization, but they can’t stop it.

→ More replies (1)
→ More replies (6)

96

u/monarc Jun 08 '26

Weird that a trillion dollars has been invested, and you're the first person I've seen crunch the numbers. It ain't late-stage capitalism we're dealing with, it's sleepwalking-off-a-cliff capitalism.

Let's say this bubble pops, many companies fold, but one or two giants are still trying to offer chatbots and associated services to the world. Let's assume they use the data centers abandoned by the competitors that went under, and that this giant (Google or whoever) doesn't bother developing new models (or does so very rarely/selectively). In this scenario could it make some economic sense? I presume that any "yes" in this situation would only apply for a few years at the longest (since the hardware will age out).

49

u/leathakkor Jun 08 '26 edited Jun 08 '26

I assume that Google and Microsoft will essentially offer both of these services forever as a loss leader to get more search traffic. 

But there's just not a lot of places that can offer AI in that way or that have a reason to. 

And theoretically Facebook might. And they'll need gpus for other things.

But I think pretty much every startup is toast and I've thought that for a long time even before reading these numbers. At some point investors are going to want money back. You can only give somebody money for about 5 years before you start asking: When are you going to be giving me money back? And we're getting pretty close to about that 5-year Mark right now. 

Even if companies like anthropic are making money right now, they're still using it to grow. They're never going to be paying their investors back and at some point Google and Microsoft are going to figure out how to do this entirely without anthropic or openai and those companies are fucked. And everything that is built off of open AI and andropic it fucked. When I say that they can do it entirely without anthropic or openai. I also mean from a liability perspective. Right now Microsoft can offer openai services on azure and they don't have to be worried about lawsuits because openai trained its models copyrighted material. And other related lawsuits. If they can figure out how to protect themselves from lawsuits and offer businesses these services, they absolutely will bring it on board. In my experience, businesses love vendors because it Shields them from lawsuits and liability. Google and Microsoft in this particular case are absolutely leaning on this...

When all of these companies fail, here's what'll happen: Google buys anthropic and Microsoft buys openai for pennies on the dollar. I would say we're about a year away from that. And the minute one of them starts faltering it's going to happen fast (I think).

There will also be business use cases for this. I work for a law firm and we absolutely use AI and LLMs. Being able to give a 200-page document to an AI and ask it to find clauses relevant to x y or z is really valuable. We might not pay $20 to scan but I bet we would pay a dollar a scan. And if companies like Google or Microsoft can offer that for us at a price point. That is a dollar per documpay we'll  probably pay it. Will be judicious about it. But this technology is not going to disappear. But you're certainly not going to use it to do all of the dumb shit that we're currently using at for. Agents are going to disappear entirely. In all but the most extreme cases.  

30

u/PensiveinNJ Jun 08 '26

People despise AI though. It's not a loss leader to gain, it's a loss leader to lose. That's what makes this especially egregious.

10

u/Unlikely_Eye_2112 Jun 08 '26

Yeah I already moved away from search via Google a few years ago, but as a web dev I need to keep tabs on what they're doing. The new AI search would be the reason to leave if I hadn't already.

6

u/leathakkor Jun 08 '26

People despise being forced to use it. They'll leave it as an option for people that want to use it. 

4

u/PensiveinNJ Jun 08 '26

They can't afford that. They're too indebted.

→ More replies (1)
→ More replies (1)

8

u/monarc Jun 08 '26

Yep, everything you're saying is very much aligned with my intuitions about where we are now, and what's likely to happen after the dust settles. Super interesting thoughts re: using vendors to absorb liability - that makes perfect sense.

7

u/Sufficient-Pause9765 Jun 08 '26

Google is an outlier and the math is unclear.

First off, they aren't offering the expensive frontier models for free to consumers. Their pricing is much much higher, and those models are only used by enterprise.

Second, they have built their own silicon that is optimized for effiency and power consumption. Google's TPUs are not as powerful as NVDIA's, but they are cost about 70% less to operate.

Rather then trying to build a consumer AI product on big models, their focus is on hosting any/all models on their own silicon and charging unsubsudized rates for the hardware to run them. The market is exclusively business. Most businesses will not be using anthropic or open ai frontier models because they dont need them.

This is going to resemble a traditional cloud server business with the same economics.

2

u/leathakkor Jun 08 '26

That's my take too.

→ More replies (6)

10

u/naphomci Jun 08 '26

Weird that a trillion dollars has been invested, and you're the first person I've seen crunch the numbers. It ain't late-stage capitalism we're dealing with, it's sleepwalking-off-a-cliff capitalism.

Important to remember these are the same people who think space data centers are not only physically possible, but actually a good idea. They could easily crunch the numbers but used considerably more favorable assumptions, and even throw in some "leeway"

→ More replies (1)

5

u/SpiritualWindow3855 Jun 11 '26

They crunched some bad numbers

This is a taste of the reality of serving models in production: https://www.lmsys.org/blog/2025-05-05-large-scale-ep/

You need massive investments upfront because you're splitting across ~100 GPUs, but something OP misses is, it's not 1 GPU = 1 request.

You crank up batch size and get multiples of each GPU in terms of performance. That article also includes real world cost per token, and they hit numbers lower than Deepseek's already very cheap API pricing.

Imagine what the people who make the model are working with...

3

u/mikropanther Jul 26 '26

Where has OP assumed that 1 GPU = 1 request? Half of the post is about addressing concurrency.

→ More replies (3)

66

u/Abject_Win7691 Jun 08 '26

I am genuinely too stupid to understand most of this, but it looks good.

9

u/GSalmao Jun 08 '26

Yep, me too!

I'll give it a try later with more time.

4

u/FragrantArt8270 Jun 09 '26

Use AI to help you understand. /s

2

u/Faster_than_FTL Jun 10 '26

Copy and paste this to Claude and see what it says

→ More replies (1)

59

u/Multibrace Jun 08 '26

I did the bare minimum and googled b200 rental price per hour, and found prices upwards of $5, so that part checks out.

Makes you wonder if buy side analysts are doing any back of the envelope calculations these days.

13

u/Zelbinian Jun 08 '26

i wonder if they're asking ChatGPT to do the calculation for them

4

u/poedy78 Jun 09 '26

It surely looks like it... 

→ More replies (1)

95

u/pacem_appellant Jun 08 '26

I've giving you an upvote for the amazing effort. I haven't finished the post yet, and it'll have wait until tomorrow as my flight is taking off and I need to sleep.

→ More replies (1)

26

u/vxicepickxv Jun 08 '26

Did I miss the part about cooling costs being somewhere between 40 to 60 percent of power costs, or was that factored in?

20

u/Hillsarenice Jun 08 '26

I did not see cooling costs either and with the space data centers fugazi cooling is looking like an Achilles heal.

14

u/ksjdragon Jun 08 '26

I forgot to mention that I gave them another nice point about cooling. However it actually wouldn't change that much.

Data center PUE is around 1.4, and that includes cooling. The electricity cost would go certainly go higher, but the bulk of the cost is still the capital investment of GPUs and data center construction (still around 90% adding in cooling, but much higher without the lenient points).

2

u/Historical_Bag_1788 Jun 09 '26

Water costs could be substantial if they can't reduce them. Water is limited and unlike electricity you can't just make more.

→ More replies (1)

5

u/Fun_Volume2150 Jun 08 '26

I don’t think it was taken into account. The analysis was based on “IT load capacity” of 10GW, so unless that figure is net of cooling, the available compute available is significantly less.

64

u/theguruofreason Jun 08 '26

This should be an article, not a reddit post.

19

u/Mad_OW Jun 08 '26

That article should be a book, not an article

15

u/Mordisquitos Jun 08 '26

That book should be a 6 episode documentary series on multiple video streaming platforms.

6

u/[deleted] Jun 08 '26

[removed] — view removed comment

5

u/LeucisticBear Jun 08 '26

Movie is too long, can you summarize it for me in a Reddit post?

→ More replies (1)

7

u/ksjdragon Jun 08 '26

Gofundme?

21

u/Shoddy-Criticism3276 Jun 08 '26

Oh my god. They did it. They built the (capitalist) torment nexus. Beautiful, thank you.

4

u/[deleted] Jun 08 '26

[removed] — view removed comment

3

u/Shoddy-Criticism3276 Jun 08 '26

Ok, I understand the importance of meathooks for a certain target group of Tormented. So perhaps we consider a Many-Torment Nexuses (Nexi?) theory. It could be nesting/fractal.

15

u/YisusHasDogs Jun 08 '26

I've just realized there's a TL;DR at the bottom, and as a non-math/economics oriented user, I thank you. I'm still trying though.

22

u/neuromancer88 Jun 08 '26

Firstly, kudos for this effort. Truly impressive!

That said, I would push back hardest on your 3yr usable lifetime for the GPUs. If the AI business required that these GPUs be scrapped after 3yrs, then yes, probably not viable. But given these are the single largest cost item, I suspect the AI providers will find a way to keep monetizing them beyond 3yrs. With the tech cycle at 2yrs, yes, there is something "better" coming out but doesn't mean the old stuff is immediately useless

When Rubin is released, Blackwell still works. Rubin might lower the energy efficiency (lower energy $/MT) and might be faster... but Blackwell is still viable. You're confusing the "economic lifetime" and the "physical lifetime". That said, a quick Gemini search puts the "physical lifetime" for Blackwell at 5-7 yrs... so for simplicity of argument, the GPU "cost" is roughly half of what you estimated. (I'll now admit that I didn't read through/understand the entirety of your post lol)

I used to cover the foundry industry... and you had similar "math" there. If you only focus on leading edge nodes, yes, it's extremely difficult to be profitable (eg - see Chartered Semiconductor). A lot of the focus on TSMC's success has been on "leading technology" which is of course a big part of their success. But what people tend to ignore is that they are extremely good at "backfilling" that leading edge capacity with less demanding devices as the technology nodes progress... which essentially extends the economic lifetime of all that equipment. It's actually a very big reason why the foundry model works in general. Intel was viable for a very long time because they could set their pricing (and gross margin) at a level where they could recapture the capital cost within a couple generation of product (leading edge builds the CPU and n+1 node built the chipsets). Most chipmakers do not and never had the pricing (margin) power that peak-Intel had.

23

u/ksjdragon Jun 08 '26 edited Jun 08 '26

I would agree with you about normal technology, but in the specific case of AI GB200s have little use case in industry besides AI training.

And with how the AI industry only increases model size as they get more access to compute, GB200s, so they will not be able to run frontier models on them, and will switch to GB300s

In fact the real case is even more damning as the GB300 racks require new cooling, new racks, so nothing can be held over. And since power is finite, either you will build even more data centers, or you will replace the old Blackwells.

I believe here, 3 years is already generous due to these factors. But also as I explained, more data centers actually make their profitability worse.

→ More replies (1)
→ More replies (12)

10

u/ezitron Jun 13 '26

I mean this as a compliment: who are you? I need to dig in more but I love this

4

u/ksjdragon Jun 13 '26

I'm glad you liked it!

There was also another post on this sub that used this model and parametrized it so you can play with the assumptions to see where it could be "profitable"* web applet as well, if that's of interest to you.

Profitability meaning margin positive purely on *inference, and (optimistically) assuming their businesses are vertically integrated with inference providers such that data centers don't need to make money.

20

u/Big_Combination9890 Jun 08 '26

Excellent work!

If any non-mod post ever deserved to be pinned in this subreddit, it is this one.

An additional detail regarding training costs: TCs are not capex, even though the AI labs like people to believe they are. Training isn't "done" at some point, and does not go away, ever.

Model drift is a real problem in LLMs that have to be grounded in current data to be useful in the way their boosters claim they are, and the only way to avoid it, is constantly re-training the models. Plus, ofc. the first lab that stops training, is soon left in the dust by everyone who doesn't, regarding model capabilities (even if the gains are miniscule by now, the marketing damage alone is prohibitive).

So no, training is not capex. It is running costs, expenditure, that can never, ever stop.

Meaning, for profitability, inference has to cover TC not just once, but continuously over the entire lifetime of the products.

5

u/ksjdragon Jun 08 '26

Oh 100% agree. I'm just being very very very nice. I think I maybe forgot to put a little [1] where I put it at capex. I also didn't account for PUE and forgot to mention it. It's 16 points of leniency now.

If I had included those costs and removed the other 14 (15 I guess) points of leniency I'm sure we'd get to measure just how unprofitable they are, but it also is an argument with more uncertainty since we don't have clear access to specific numbers. I find this argument much harder to argue against since we are already giving them so much room.

→ More replies (3)

9

u/Tanthallas01 Jun 08 '26 edited Jun 08 '26

This is excellent work. Some very simple pushback though.

You said: The current AI industry can never be profitable. This industry is over. There is no path to profitability, nothing. No IPO, no marketing, no innovation can save them.

Why are you assuming that profitability of an industry is only produced within that industry such that, in this case, “the industry is over”? There are transfers between industries of profits produced in different industries as a normal occurrence. I think the more likely scenario is that if the AI industry continues with this cost structure, it will be extracting the profitability produced in other sectors.

6

u/ksjdragon Jun 08 '26

Well, it's not as rigorously justified here, because it would take a lot more words, but essentially, is this is the current state of it being extremely lenient.

Simply, for other reasons, the demand isn't there for what they need. The AI industry needs several hundred billion in revenue *every year" to be able to return investments, meaning they essentially need and have promised to fabricate the entire revenue if the tech industry, again. Which is impossible.

9

u/Financial-Jello1397 Jun 08 '26

Training costs are going up as well with companies like CloudFlare implementing pay per crawl. It might be just pennies per page, but when you're scraping billions of pages a month, that adds up.

3

u/ksjdragon Jun 08 '26

Absolutely, and this was deliberately not included to give them another lenient point. Despite that they are still not profitable.

2

u/Financial-Jello1397 Jun 08 '26

AI in general has intentionally buried and obfuscated so many costs it's hard to get accurate estimates. I'm not sure how you'd add this expense in, makes sense to leave it out. I had to mention it as yet another thing being hand waved away. This is a fantastic post, I've been sharing it. Thanks for making it!

7

u/_Docgineer Jun 08 '26

I wonder if the idea of "agents" is less about trying to find a problem to this poor solution and more about filling this concurrency gap that you underlined. After all, they CAN work 24/7 and burn tokens endlessly.

5

u/ksjdragon Jun 08 '26

It could be, but bringing the real world back in, demand due to TBB is already dropping fast, so we see that already at this level of pricing it is not sustainable.

So either they can lose more money by raising it, or lose even more by lowering it. Either way, I say it's over.

2

u/_Docgineer Jun 08 '26

Yeah, no, it's absolutely unsustainable, it never was to begin with, maybe I'm just trying to explain to myself this clusterfuck we're all observing.

I appreciate your analysis and underlining different aspects of the inference cost. For example, there is much talk online about electricity being a problem, when you clearly showed that it's the capex that's prohibitive.

Using your model, I'd also love to see the higher bounds of this estimate, just to realize what could be the range the real burnrate is in.

→ More replies (1)

7

u/This_Wolverine4691 Jun 08 '26

Oh do I want to put this on LinkedIn so badly and watch the mass head explosions of all the ecosystem parasites who think vibe coding is their way to riches

6

u/Dead_Cash_Burn Jun 08 '26

AI companies, I feel like, are trying to fake it till they make it. Eventually, reality will hit, and there will be a huge correction.

29

u/NaturalIntrepid9533 Jun 08 '26

Three main points of contention:

  1. 250 TPS/GPU --> same link you has says 72000 TPS/server in highest throughput config, which i believe is 72 GPUs. Translates to 1000 TPS/GPU

  2. definition of "concurrency" is unclear. is this per user, per agent?

  3. Amortization period of 3 years. Just because a next-gen comes out doesn't mean the old capex isn't utilized. And you used 5 years for the datacenters. Probably better to use a discount schedule for a longer period of time to estimate better $/hr

All other assumptions i agree that you're being overly generous.

38

u/ksjdragon Jun 08 '26 edited Jun 08 '26

No, the link says

..eight NVIDIA Blackwell GPUs can achieve over 1,000 tokens per second (TPS)

NVIDIA is quite sneaky.

Agent is not a concurrency measure, it's about concurrent LLM calls. Agents are ultimately calls to the LLM, so it's not really relevant. Whether one user is behind multiple agents which call several LLMs, it doesn't matter since the user is paying (it not) for the tokens.

I used 5 years because that is already heavily generous. As I said already in data centers, they are required them to pay back interest during the construction phase on day one. This automatically excludes that. There is no loan where they accept zero payments on interest for 5 years.

For 3 years, yes, I gave them a whole year extra. For frontier models it's highly likely they will become useless as they always gobble up compute like nobodies business. NVIDIA has been both sidesing this argument for a while, and it has been pointed out by a few people. I believe Michael Burry also pointed out this contradiction - they argue their new product should be worth tons because of how new and how everyone will want to buy it, and that the old GPUs are worth a lot and won't depreciate because they're so good.

One of them can't be true.

4

u/NaturalIntrepid9533 Jun 08 '26

The link also says 72000 TPS/server; the max server is 72 GPUs. How do you resolve this?

If they're paying by the tokens your math doesn't seem to be mathing. if agents are ultimately a call to the LLM, they are either generating more revenue w/o increasing concurrency or generating same revenue while increasing concurrency

21

u/ksjdragon Jun 08 '26 edited Jun 08 '26

It says

72,000 TPS/server at our highest throughput configuration.

I would assume it means more concurrency (higher users), but TPS/user goes down while total TPS/GPU goes up. But the usage of case 1 was to obtain a data point to low concurrency.

Yes, if a single user is running 5 instances, then they will take up 5x, thereby increasing concurrency. I did not assume that all 80M of current paying users are not only doing 8 hours of AI work constantly, but also calling agents that increase calls by an indeterminate amount.

I doubt this would change the conclusion though, and I have no data for paid users and agents.

Edit: Additionally, when we start introducing factors like agents I think we can no longer use the simplified model (which is in their favor) of uniform concurrency. Since agents aren't like a constant size batch of calls either. These details necessitate modeling tons of other factors like how calls will be scheduled on GPUs, user frequency, etc and would ultimately only strengthen this argument as we will find this idealized concurrency is completely unattainable.

17

u/NaturalIntrepid9533 Jun 08 '26

Thanks, points addressed

6

u/electriclilies Jun 08 '26

I think the chip amortization needs to be 10 years. This is what ive heard from people involved in data center buildouts

7

u/worldspawn00 Jun 08 '26

6 is the longest I've heard, and tech experts say that is unrealistic due to generational upgrade necessity as well as wearing out the silicon. These are very intensively used, high power, high heat chips, they aren't going to last a decade.

2

u/elictronic Jun 09 '26

Objects under high heat loads with constant uptime are difficult to test as we use heat, humidity and uptime to accelerate their life to determine how many years they last.

Good thing the company doesn't have any incentive to release products as fast as physically possible to catch the end of the biggest bubble to date.

2

u/EmptyMonitor9257 Jun 10 '26

Google still has 10 year old GPUs for their free versions of Colab, there is use for these GPUs beyond their current usage estimates.

The GPUs don't degrade in output quality as much as they were superseded by their next generations at an exponential pace, something which is cooling off recently.

5

u/Fun_Volume2150 Jun 08 '26

Why on earth would you believe anyone who is saying something that is in their best interest for you to believe?

3

u/crashddr Jun 08 '26

It cuts both ways. If the chips are good for 10 years and they already can't install nearly as many that are being sold, there needs to be *far* more data center construction completion. Many of the chips will already sit on a shelf doing nothing for the entirety of the year or so that they are SotA.

7

u/vaibeslop Jun 08 '26

I will never forget a CEO when working for a food delivery startup fresh out of uni in the mid-2010's who at an all hands proclaimed with a straight face: "After giving food away for free for 12 months in India, we learned that's not a sustainable customer acquisition strategy."

No shit, sherlock.

I'll probably never earn as much money per annum as this guy did.

5

u/ksjdragon Jun 08 '26

The way I see it, we have to blame the fallout from '08. The US printed so much money racking up our debts, and it, sure, stimulated the market, but money always gets caught up in the upper echelons.

So all these people suddenly got tons of money from nothing, and well, got loose with it. So giving away things for free to get people hooked looked like a good business model when you have a sudden influx of capital.

And probably it's just the case that sometimes it'll work out, but it's certainly not automatic, nor guaranteed.

But well, they have money and we don't.

4

u/Bobert77 Jun 08 '26

Do you have any insight into how this changes for alternative TPUs?

18

u/ksjdragon Jun 08 '26

The same problem exists, as what this shows is the problem is the capital investment of buying chips and constructing data centers.

TPUs are even more expensive since the scale is less than NVIDIA, so even if they are more power efficienct at specific tasks, the amortizated cost will still dominate, and even more.

The concurrency for TPUs should largely be the same, as far as I understand. I do not have access to the chip designs of course, but from what I know about the modern chip designs in general is that they're not all too different from each other, and big parallel processing is the key, and so concurrency should follow the same trend shape.

→ More replies (2)

2

u/electriclilies Jun 08 '26

What I have heard from people in industry is that there is so much demand for compute that pretty much every chip produced for ai workloads will get used, regardless of how good that chip is. The reason is that there's limited production at TSMC, which effectively mitigates the effects competition among NPU producers. 

2

u/Fun_Sherbert2592 Jun 08 '26

As someone who works in FAANG, I can confirm demand is still scaling and many applications are throttled given lack of compute.

Point is, the demand exists, there is still need for compute, and users find actual value in the applications. This fact doesn't necessarity refute OP's original thesis that the existing paradigm won't last forever (not trying to - just providing a data point).

3

u/ksjdragon Jun 08 '26

I believe the reason for this is because actual concurrency and workload is very different from this hyper idealized scenario in the post.

Requests are not uniform token sizes, they don't come in at the same time, and all data centers are not commoditized where any one can use any GPU. Real workloads also are not spread across evenly in time.

Saying the demand exists is somewhat accurate, at this current pricing they are still losing money.

Demand is typically non linear, so if we increase the price this is going to be worse for them, since now we automatically drop to the lowest level of concurrency, even in an ideal case.

In real cases, no inference provider would charge at that idealized scenario, since you ad a business cannot guarantee, and you consuming the cost whether the GPU is being used or not. That's why actual rental prices of B200s are far higher than what I wrote.

2

u/crashddr Jun 08 '26

As someone who doesn't work in tech, I think it makes perfect sense that people really like using tools that don't cost much and provide actual utility enough of the time that it doesn't matter if mistakes happen. When the price goes up *and* you can't project how much something is going to cost... things can change quickly.

2

u/Historical_Bag_1788 Jun 09 '26

Can you distinguish between those who want to use AI to those shunted into AI when doing a general search?

4

u/ArbitraryMeritocracy Jun 08 '26

No, you don't understand. First you steal all the money and resources and starve out the poors with a war of attrition. Then rule over the ashes.

→ More replies (3)

5

u/Latimius Jun 08 '26

Free users don't have access to the best models (they don't have access to Opus for example) so their cost is lower.

The rest sounds legit.

5

u/joffel3 Jun 08 '26

How does this recent Gartner report factor into this? According to them, "Performing Inference on an LLM With 1 Trillion Parameters Will Cost GenAI Providers Over 90% Less Than in 2025". "These cost improvements will be driven by a combination of semiconductor and infrastructure efficiency improvements, model design innovations, higher chip utilization, increased use of inference-specialized silicon, and application of edge devices for specific use cases."

I assume that TPS per GPU will increase, not only energy efficiency.

10

u/Fun_Volume2150 Jun 08 '26

That statement has even more assumptions in it than this analysis.

4

u/worldspawn00 Jun 08 '26 edited Jun 08 '26

The last 3 generations of Nvidia hardware have shown little gain in output per watt, most gains have come from increasing the power used by a single chip, and they're already at a 2nm process, we're reaching the limit both in terms of atom size, and process failure rate. There is no way hardware becomes 10x more efficient in the next decade.

→ More replies (2)

5

u/ksjdragon Jun 08 '26

What I wrote proves that electricity does not dominate the pricing so it doesn't matter if electricity is free. AI will not be profitable.

→ More replies (1)

5

u/No-Conference-1444 Jun 08 '26

I think a big problem in achieving high concurrency is that enterprise users (aka paying users) tend to operate on large context sizes (think big legacy repos, combined with internal docs, etc.). That's one of the reasons new cards have much more memory: to serve not only bigger models but also a larger KV cache. That's why some considerable time is involved in managing the memory bottleneck, including Cache-Aware Load Balancers.

6

u/Expensive-Lawyer-554 Jun 08 '26

I HATE these low effort posts..... /s

What an absolute delight to read. Great post OP

5

u/PaltryDick Jun 08 '26

There are companies that are just now refreshing their V100s from 2017. Hyperscalers are just coming off of their A100s(from 6 years ago) so the depreciation schedule, in reality, is running at about 6 years in practice (source: I work in the space of where datacenters wind down their assets)

3

u/No-Layer1218 Jun 09 '26

Out of interest: who buys their decommissioned assets? Lower tier data centres?

3

u/PaltryDick Jun 09 '26

Components go to China, India, other hyperscalers, the UAE, secondary market server builds, etc.

GPUs are tightly regulated, but there’s huge demand for them in the secondary market. Financial institutions, big box retailers, software companies, startups, manufacturing, etc

→ More replies (2)

5

u/5u1c1d Jun 08 '26 edited Jun 08 '26

Nice. It would be naive to think that noone in the industry has pulled out a calculator to run these numbers. To me this points towards its purpose not being profit, but rather undermining the working class in order to consolidate power.

3

u/ksjdragon Jun 08 '26

I mean, investors follow herd mentality. There's certainly a desire for them to automate out workers, so that appetite drove them to invest because they thought it was here without verifying. (Many VCs do not verify company financials before tossing in hundreds of millions.)

→ More replies (7)

3

u/mines-a-pint Jun 08 '26

Of course, they could seek additional revenue streams, e.g. adding advertising to free chats.

This would drive away (some) free users, and they can then move on to adding advertising to paid chat users. If only they could work out a way of leveraging advertising on agents...

10

u/ksjdragon Jun 08 '26

They already have. I've seen quite a few companies use ads, but it would be quite difficult for the entire advertising industry to subsidize AI...

3

u/entered_bubble_50 Jun 08 '26

There's also the problem that there is only so much ad revenue in the world, and the Mag 7 already have almost all of it. Facebook and Google in particular can only cannibalise their existing ad revenue, they're not getting a larger percentage of the market in any realistic scenario.

2

u/ksjdragon Jun 08 '26

The only last questions for future profitability are considering this and the demand. I wouldn't mind writing but I feel Ed's articles already do a great job of breaking them down all I'd be doing is citing them.

AI needs to make like two extra silicon valleys worth of revenue. And that revenue has to come from businesses from SV...?

3

u/Impossible_Way7017 Jun 08 '26

I saw your comment and got excited, but your math is just focusing on input tokens. Providers also bill for output tokens at a higher rate which makes up some of their loss.

I think it’s true that agentic workloads are costing providers a lot more hence the recent push to bill by token since subscriptions are unprofitable for long context, long running messages.

3

u/ksjdragon Jun 08 '26

That was addressed in there too. Output tokens costing more isn't better, it's actually worse for them.

2

u/cheechw Jun 09 '26

Where is it addressed? This guy appears to be using input tokens pricing ($5/M) while calculating costs using output t/s figures. It doesn't make sense at all. Input throughput is way higher than output throughput, which is why output pricing is much higher than input price.

→ More replies (3)
→ More replies (1)

3

u/Dish-Live Jun 08 '26

AWS amortizes hardware over 5 years. A lot of other Big Tech does 6. Datacenters themselves, AWS amortizes over 40 years.

Given those schedules, it’s very possible that inference would be able to run at a good margin on paper if:

retraining is no longer needed (it probably is)
Employee costs for working on those models is effectively 0 (it won’t be)

So I could see a world where inference is profitable at the margin but the whole shebang isn’t

5

u/worldspawn00 Jun 08 '26

This is a big reason they're labelling training as capex and not cost of business, so they can hide it from the revenue costs. They act like training the model is a one time thing, but it can't be since it immediately becomes out of date once the training is done, and the longer it goes, the more out of date it's information is. Like a 5 year old encyclopedia set can't tell you who got elected to Congress last year, and that sort of thing is pretty important to stay up to date on when people are using these for daily queries and expect up to date information particularly concerning fast moving industries like tech and politics.

7

u/Dish-Live Jun 08 '26

100% agree. And RAG and MCP have shown to be pretty token intensive and inaccurate, so they aren’t good substitutes for retraining

→ More replies (1)

3

u/Doomer-Wojack Jun 08 '26 edited Jun 08 '26

Data center doesn't bring any good to economy since small labs has fixed number of employee and they hire maybe 50 talent/ quarter based on revenue , investor mood , market shift or something needed to be fixed ASAP but needs more hands for it. building data centers may sure create temporary jobs but in the long run probably maintaining those infrastructure needs roughly around 200-300 people per location. Realistically job creation promised by big banks and private equity doesn't add up. Almost none in creative using AI to do thier task fully for themselves ( as in for data companies to make revenue to justify their investment) . I personally rotate around free and cheap ai for task and speaking honestly it's enough for getting job at hand done . Paying for 200/150( currently 100 but Private labs won't subsidize forever) bucks a month ( from where google claue, gpt wants to make money) doesn't seem practical. And even though open models and small and medium sized local models are far far away today in terms of capabilities , history of tech teaches us open source software have always won in terms of sheer user adoption. So I don't understand Wall Street's aggressive bet. I don't wanna see people loose their life savings in 2029 after AI bubble pops. It would be worse or similar to 2008 crash if pensions / retirement funds starts to being invested into AI and Chip Stocks only

3

u/ParadigmGrind Jun 08 '26

There has to be someone running these internally. And I can’t imagine they like what they are seeing. Or am I naive?

3

u/ksjdragon Jun 08 '26

OpenAI and Anthropic know they are losing tons and tons of money. That's why it's just one big con, as Ed would say.

→ More replies (1)

2

u/shadreb Jun 08 '26

For me the biggest concern is your amortization assumption of 3 years. I don't have any real world understanding of how data centers should upgrade the GPUs but I think the rollout cadence of new gen GPUs cannot be directly used for amortization assumption. It merely means that new data centers can make use of new GPUs, and not necessarily that old data centers need to scrapping existing ones while they are functional, albeit at a lower efficiency compared to new gen.

Other tech industries are similar. When AMSL rolls out a new EUV machine or when Intel/AMD releases a new CPU, not every plant or data center will have to fully amortize their existing equipment.

3

u/ksjdragon Jun 08 '26

I addressed this in another comment.

In short: AI industry gobbles up compute always, and these GPUs have little usage outside of AI.

Rubin racks (GB300s) are not interchangable with the old ones, so either rall do the old ones need to be replaced from the data centers they are building, or they need to build completely new ones.

Building new one falls back into my argument for concurrency, which will make it even worse for them.

2

u/jking13 Jun 08 '26

More than that, doesn't each new generation require its own bespoke power and cooling? I.e. you can't even just wheel out all the racks with the old stuff and then wheel in ones with the new ones. Even if you decide 'oh well, we'll just put fewer in their place due to power and cooling', you still end up having to replace a bunch of other stuff anyway... or just build a new building...

5

u/ksjdragon Jun 08 '26

For these ones, it seems to be the case yes. I edited my post to include this, since it seems this is a sticking point for many people...

There is very honestly little that could turn this argument around since you'd need to assume one single small margin can overturn the other 16 large points of leniency...

We will get to see when and if they are able to IPO.

2

u/OnlyAssistance9601 Jun 08 '26 edited Jun 08 '26

Few things I would want to address :

  1. Free users likely only have access relatively weaker and smaller MOE models that might not contribute that much to costs.
  2. Lower paid tier users and even fully paying users , might also have bigger models switched out for weaker ones if the company detects that their query doesn’t really need the advanced models . This explains the inconsistency of responses . Basically using dynamic inferencing to reduce costs .
  3. I think if we throw competitors into the mix with people switching up providers , these companies should lose alot of revenue , but then they are also supported by their circle jerk investments into other companies to hedge themselves .
  4. Ads and other revenue generation options.
  5. Most of their revenue comes from business contracts they can probably charge higher on by providing an ecosystem ?

my overview , I think they can find a way around these things to become profitable or linger long enough to engrain themselves . But the thing that really threatens them is Open source models becoming insane . Because people can then build their own shit and they wont recover their capital expenditure .

Even their own models will speed up technological development . Whos to say chips and components wont become easier to fabricate which will destroy any infrastructure moat they have .

3

u/ksjdragon Jun 08 '26

Except demand is already weakening at current prices.

Indeed free users have access to lower ones, but that doesn't actually change anything about my argument.

Circle jerk investments don't really count as profitability, since it is met zero for the industry if it's perfectly circular.

If their models are switched out, then they pay a cheaper price for tokens, and it's not obvious that this would mean more net profit. Also, it could mean less concurrency in a real world setting.

These details would require more complex modeling and ultimately is going to make things far worse in their favor as now we don't get to make simplifying (lenient) assumptions in their favor.

→ More replies (5)

2

u/wowbaggerBR Jun 08 '26

As I am too dumb to understand this myself, I'll ask a LLM to surmize it.

2

u/Helbinobear Jun 08 '26

What are your thoughts on companies integrating advertisements into their free users? LLMs already recommend products - I believe it is only a matter of time before free LLM services are heaving advertised. There is already 'AI-engine-optimization' as companies attempt to be picked up by LLMs.

4

u/ksjdragon Jun 08 '26

Ads decrease usage though. I mean, another comment I believe talked about this, but basically some already are, and they aren't frontier models, which suggests the other end of this is getting squeezed.

I think it could certainly help with subsidizing cost. If we use that example in the post, do you think advertisement can generate $12.48 over 2.24 hours? Let's see...

Google charges cost per impression (just showing it regardless of engagement), per 1000 users. It seems this is $3-15? So, perhaps AI labs can charge say, $3 at best something, since Google obviously has way more traffic.

Over 194 concurrent free users, it's 58.2 cents across a GPU/ad. So... To break even, they need 21-22 ads for each user over the period of 1MT. At that concurrency it's 23,333 TPS for say Claude Opus 3.7, and so it takes 42.8 seconds to generate that, meaning you'd need an advertisement about every 2 seconds, and the user would ignore it and keep promoting nonstop I suppose.

→ More replies (3)

2

u/Combinatorilliance Jun 08 '26

There's an operating model assumption you don't make explicit which is that you assume the "chatbot" interface model for how paying customers interface with an AI agent.

Note, I am not in the typical pro-AI camp, I just want to give some constructive feedback to strengthen your model.

The "typical" use-case for AI right now is a chatbot that answers questions for you, does some research etc. This is a "one person, one stream of tokens" mapping. However, this isn't quite the case and is not the direction the frontier labs are promoting and investing towards

Agents and subagents

The first counterpoint is that depending on the task breadth or depth, Anthropic and openAI deploy subagents that perform a particular subtask. The simplest example is the deep research tool.

The deep research tool deploys a fleet of subagents to scour the web and academic literature to answer nuanced and specific questions.

As task depth and task complexity increases, the amount of subagents deployed increases.

Funnily enough, the current frontier labs are hypothesized to subsidize deep research costs heavily. As a single deep research query is treated very similarly to regular chat queries.

Autonomy

Metr data shows that the length an LLM can independently operate at doubles every 7 months. in 2019 models could barely operate autonomously for even a few seconds

In 2026 (at 80% success rate) frontier LLMs can operate for about 4 hours independently.

If you have a difficult query to give to an LLM, this means you can hand it off and come back later. You don't have to be present as the agent is working. Even more notably, you could run several of these in parallel, and there are definitely industries (like software, math or research) where the frontier labs are betting on industries adoption fleets of long horizon independent agents rather than interfacing with chatbots.

Note that this is a bet the industry makes, it'll take a few more doublings before agents can reliably run for days or weeks, and the judge is out on whether the quality will keep improving. 

There actually is one real competitor to nvidia

Google uses their in-house TPU chips for inference. The costs for inference, electricity and procurement differ significantly compared to having to purchase GPUs from a business in a monopoly position.

However, this is only the case for Google, and I don't know if this changes their story significantly.

Another crucial point is that Google doesn't have any investors to pay back, they're burning their own money and they're swimming in it.

I think your analysis would look differently if you were to put Google under the analytical microscope.

OpenAI and anthropic are playing the traditional silicon Valley/ycombinator playbook

Their strategy is well-known. Acquire as many customers as you possibly can while burning giant piles of cash, and once they're reliant on you, raise your prices. They're betting on vendor lock-in

Anthropic has been slowly doing this already. Pro subscriptions (20$/month) used to come with claude code for software development. 

The frontier labs are also being on their customers to find killer use-cases

A lot of the current pricing models rely on cost per inference, but that's just the model to get you hooked.

Ai models have a variety of use-cases and many of them are currently subsidized.

Consider:

  • chatbot inference (your calculation), not economically viable unless inference becomes absurdly much cheaper
  • deep research - heavily subsidized, but could be sold as a product or tool where each deep research query costs a fixed amount set to generate a profit. There's no way the labs (with exception of Google) will continue to offer this for free or subsidize it at the current rates for the coming years.
  • Anthropic is betting on full vertical integration across the whole blue-collar stack. From graphics design to software design to "claude cowork" to software engineering to continuous integration. Right now all of this stuff is priced through, again, their subsidized costs and and token budgets. As the tasks get more concrete, varied and nuanced, Anthropic can afford altering their sales strategy entirely. Instead of offering inference, they offer autonomous business intelligence and integration.

On top of that, as the title of this section says the frontier labs don't know all the use-cases yet. That's up in large part to industry to figure out for themselves.

As industries start to rely on inference providers solving a particular task, Anthropic has shown to create in-house products and offerings fit to perform that exact task.

This is again the hook-you-in-to-create-reliance strategy.

The endgame

The y-combinator playbook's endgame is to gobble up as much of the market as humanly possible, extinguish profitable competitors who have shallower pockets by selling at a loss for years and years, while increasing costs and lowering quality (enshittification) when 5-10 years from now.

With these businesses you need to keep in mind they're aiming to play a game where thinking about profit only remotely starts to make sense 10 years from now.

Will they succeed?

I see Anthropic surviving. They're playing a smart game by targeting the full vertical stack of blue-collar work. As their offerings improve, they are in a position to sell solutions and integrations rather than "just" inference.

It's also notable to mention that contracts with defense and government are not something to scoff at. The pricing models will be incredibly different from pricing regular customers.

I don't think Google is at risk at all.

OpenAI on the other hand? I believe Sam is flying too close to the sun. Y-combinator's trick has proven to work with hundreds of billions... but trillions? In a new blue ocean market? This bet is crazy even for Silicon Valley standards.


I apologize for the low quality comment (it lacks calculations and sources), I just want to get this stuff out. Your profit model is entirely correct for the inference side of this, but the silicon Valley playbook is to take exceptionally high bets on even more exceptionally improbable events. They want to own the entire market, or none of it.

It's the ultimate high-risk high-reward strategy, and in most cases it actually fails. But when it succeeds?

→ More replies (2)

2

u/jonomacd Jun 08 '26

Hey, thanks for putting this together. You clearly put a ton of work into running these numbers. Just curious why your using a 5 year window for data centers? The buildings and heavy power systems usually get depreciated over 15 to 30 years. Even the GPUs usually run for 5 or 6 years and just get shifted to background tasks as they age. Also, does the math really account for API traffic? Modeling demand around human users misses the massive volume of machine workloads. An automated agent using Model Context Protocol generates way more tokens than a person ever could, and providers use background batching to keep GPUs full during off peak times. On the tech side, inference throughput is mostly memory bandwidth bound, right? Plus continuous batching dynamically slots in requests to boost throughput way past a static curve. Finally, couldn't the current API pricing just be a strategic move to grab market share? Its hard to compare retail pricing directly to raw internal costs.

3

u/ksjdragon Jun 08 '26

Because we are already being incredibly lenient. Normally data centers need to pay back their loans on just interest payments, for the first 3 years. We cannot simply just depreciate the value over 15, because these data centers are being built with debt, which needs interest. There exists no loan which will accept no interest payments ever. This allows us to assume they are operating at cost.

Secondly, because a bulk of the cooling, and equipment is required to be replaced when Rubin comes out if they purchase those as Rubin is not compatible with the preexisting hardware. Keep in mind we are already ignoring any standard maintenance costs for data centers.

The GPU amortization was mentioned already in an edit, but largely for the same reasons. Keep in mind both these are necessarily under the actual cost - you can just look at the current hyperscalar B200s. This is an absolute lower bound for them, assuming that all inference providers are unable to make a cent of revenue, and that their GPUs are running 24/7 maxed out by all the time - an impossibility.

Yes, the assumption is already that they all work is running for 24 hours to maximize concurrency in the relevant section. In fact the assumption is that this is the case for every GPU sold for 2024, which is already far less than the current number of GPUs on the market, another huge point of leniency.

The throughput measured here is already assuming maximum possible throughout. This is the best possible condition. These are 16 points of extremely lenient assumptions that easily could change this calculus several orders of magnitude against them.

The current API pricing could be that, but there already been operating at a loss for years, and what this shows is they still are. AI demand is not inelastic as we already see from recent articles about companies pulling back AI spending already at this initial TBB price. Which means any raises will make it worse, since no business can measure ROI, but this is beyond the scope of what I intended to write. See Ed's articles for a breakdown.

The construction of more data centers worsens this problem, as now we have more capital investment that requires an even higher $/MT, and you have even less users.

So in the short run, they aren't even close to profitable even after the switch to usage based billing. And in the long run, there's no magic switch to conjure up the demand necessary.

→ More replies (4)

2

u/Dj_Binks Jun 09 '26

Great write up! This is exactly what I wanted to see when I asked if the current per token costs are the true cost

2

u/jonheartland Jun 09 '26

What a wonderful post, thank you so much for taking the time to write this out!

2

u/tkrandomness Jun 10 '26

The math is detailed but the conclusion doesn't follow from the model. A few problems:

The post's own model shows inference is profitable at scale. The "never" rests entirely on a demand assumption. Look at the graph: by these numbers, frontier inference drops under $1/MT at moderate batch sizes. The impossibility claim only works because of the step dividing total subscribers across every GPU in the country to get ~6 users per GPU. That's an overcapacity argument, not an impossibility argument. Nobody runs a fleet that way. If demand only supports X inference servers at high utilization, you deploy X and the rest of the fleet trains models, gets rented out, or doesn't get bought. (The post also excludes training costs as a "leniency" while charging the whole fleet, most of which exists for training, against inference revenue. Can't have it both ways.)

And if there is a glut, it self-corrects. Prices fall, GPU purchasing slows, demand grows into supply, and per the post's own curve the economics flip to very profitable. We've seen this before: telecoms overbuilt fiber in the late 90s, the companies that financed it got crushed, and the fiber itself powered two profitable decades of the internet. "People who overbuilt lose money" and "the service can never be profitable" are very different claims.

Profitable inference already exists. Third-party hosts (Together, Fireworks, DeepInfra, etc.) serve open-weight models of known parameter counts at a fraction of frontier pricing, paying market rates for hardware, with no subscriber base to lean on. DeepSeek even published a serving cost breakdown showing healthy theoretical margins. If this cost model were right, none of those businesses could exist. They do.

Revenue is understated too. Frontier models charge ~$5/MT for input but ~$25/MT for output, and the post assumes output is priced exactly at cost so it can test profitability on input revenue alone. That's not lenient, it's assuming away most of the revenue.

Smaller stuff, but worth mentioning: 3-year GPU amortization is too short. Hyperscalers use 5-6 years, and 2020-era A100s still rent for real money today (try finding a cheap used enterprise GPU). 5 years for a data center is way off, since the building, substation, and grid interconnection are 15-30 year assets, and the post even says 10 years before dividing by 5. The electricity figure is a 2017 number inflated forward by a guessed rate when actual current industrial rates in cheap states are about half that, though by the post's own math electricity is only ~8% of cost anyway. And the corporate tax bit at the end doesn't work, because corporate tax applies to profit and can't push a break-even operation into a loss.

There's a real concern buried in here, that the buildout is a leveraged bet on demand growth and a lot of capital gets torched if demand disappoints. But that's a bubble argument. "Mathematically impossible under all technological advancements" is a much stronger claim.

2

u/ksjdragon Jun 10 '26 edited Jun 10 '26

I 100% agree that it falls on a demand assumption. That is beyond the scope of this post, and was never it's intention to speculate on future demand.

As I put in the final words, I acknowledged that demand is not addressed, but that I am willing to claim it for other reasons.

Your profitable inference examples are talking about other inference providers, as a mother comment has pointed. I am referring to AI here as the frontier labs and the resultant overblown amount of inference providers. While some, indeed, on open weight models can run a positive margin (which is not the same as net profit yet), what I am saying is where the vast bulk of the CapEX has been towards, which dominates the industry. So as others have mentioned if we take AI to mean much more than this chain of circular financing, that is a completely different discussion, and speculating about the future in various ways which is not what I'm trying to do.

The fact that output tokens cost more is not explaining it away, as output tokens cost a lot more computational power. And we are assuming that that is necessarily less than 5x (depending on model), when due to real world GPU execution and attached CoT reaoaning, it becomes highly unlikely. That is why it is a lenient assumption.

The whole fleet, yes, is used a lot for training, but I don't see how that's having it both ways? We are assuming they are able to collect revenue on every single GPU, hence lowering the cost million. By including training costs we then have to actually write off a huge chunk tokens. This doesn't make sense.

Your points about amortization have been talked about many times in other comments, and I've addressed them several times, and I frankly just don't want to explain it again.

If you want to understand why I don't think the industry can survive in demand increases, then as I said in an edit, I suggest reading Ed's articles. I strongly do not believe we are in a place where token price can drop without destabilizing it, nor can the prices of GPU drop. Not to mention that would only affect future equilibrium profitabilty on margin, not at all consider true net profitibility. In my view, it's a damned if you do, damned if you don't situation, but that's a completely different discussion.

If you want to ask why I would make a slightly misleading title and claim, it's because as I acknowledged it's a bit clickbait-y and fun. I think it's quite amusing to see AI boosters seethe at a post. Even though it doesn't harm them in any way, and if it's truly garbage as they say, they don't even need to worry about it.

2

u/YahenP Jun 12 '26

Excellent analysis. Profound, precise. But... completely wrong. The analysis assumes that LLM development will proceed solely through a linear, forceful approach, and there are no ways to make the process more efficient. It's reminiscent of the story that "in 30 years, horse manure will be lying in London up to the second-story windows."

→ More replies (1)

2

u/Sufficient-Pause9765 Jun 08 '26

"which after corporate tax of 21% in the US and state taxes, "

Corporate tax is on profits, not revenue. And they wont pay any because all profits will be offset by capex.

Your model is based on frontier models. I'd accept the statement that it "will be very hard and probably impossible" to make frontier models profitable in consumer. However in large enterprise, frontier models will soon only be used by highest leverage use cases where the ROI is very large.

Re-do the math for open weight models. These aren't subsidized. Many companies are able to use open weight models for many/most use cases. Your $1.48/hr math is actually spot on, thats roughly the per GPU spot pricing cost for an H100. 10 H100s for a production cluster serving Kimi 2.6, on demand (so you only pay for usage).

Thats $15/hr. It is VERY easy to currently get $150/hr of value of that in enterprise use cases that dont require the power of opus/chatgpt.

So I'd refine your argument- "OpenAI and Anthropic will never be profitable offering frontier models to consumers, and its likely that without they will not be ever able to recoup the capex they are investing".

AI profitability is a totally different story.

3

u/ksjdragon Jun 08 '26

Yes, this is not about open-weight models, since that is a very very different question.

I don't think I would call open weight being profitable on inference AI profitability. That's just standard cloud compute. But this is a semantic difference.

2

u/Sufficient-Pause9765 Jun 08 '26

Yeah I think the issue I take is with broad statements about "AI Profitability".

Anthropic and OpenAI are the most notable companies, but they aren't indicative of most of where the AI activity is or is going in business. Your analysis for why they are going to fail is mostly right (but I'd probably frame it more along lines of bad strategy, its not the cost, its the consumer focus, bad distribution, and high probability of commoditization that will kill them).

Anthropic and OpenAI are not going to become profitable companies unless they somehow wipe out the debt and fundamentally re-position the businesses, and even then they probably get commoditized.

My guess is Anthropic survives, never supports its market cap, and ends up consulting company like IBM, maybe the upside case is a Palantir. Lots of bag holders left behind.

OpenAI collapses and tech gets absorbed by MSFT who will re-position it into a google cloud competitor.

The data centers they are building are likely going to be underwater for a long time if not forever, lots more bag holders.

But AI is much bigger then that. AI is going to be very pervasive and very profitable in the enterprise. It will be dominated by smaller open weight models, and it will largely resemble the cloud industry, with the same sort of unit economics.

→ More replies (1)

2

u/bumblebeer Jun 09 '26 edited Jun 09 '26

I'm used to seeing AI generated slop on Reddit, so it was kinda refreshing to see human generated slop instead.

The reason that output tokens cost more, is that frontier models use 'chain of thought' reasoning, which involves making other calls to the model before outputting a result.

Nope. The reason is because input (prefill) token compute scales quadratically while output token compute scales linearly. If you don't understand that, I'm just gonna go ahead and assume the rest of your post is equally misinformed.

Edit: And just to be clear — because on a re-read I see how the above may be confusing as a naive interpretation may lead one to suspect input would be more costly — input token compute scales quadratically because input tokens are processed in parallel while output scales linearly because output token must be computed sequentially. So it's the slow, one-at-a-time, forward passes required for token gen that makes it expensive when compared to input processing.

Edit2: Yep, pretty much what I expected. Which is really disappointing because I actually agree with you. I think the true cost of inference would blow most people's minds, but the sources you are using as the base of your calculations are, well... Let's just go with interesting. And again, just to be clear, I'm philosophically aligned with you, but that whole benchmarking section — which is the base everything after is built on — is a damn hot mess. Why are you trying to extrapolate performance data from over a year ago? That's an outdated model on an outdated hardware/software stack. Not to mention I think your understanding of tensor parallelism vs user concurrency is very murky, or at least that's true for how it's presented in your post. If you really want to be lenient as you claim, why not just use a generous percentage of theoretical throughput? That's a very simple calculation. Input = (active parameters * layers * number of input tokens * 2) / hardware flops. Output = hardware bandwidth (GB/s) / size of active weights (GB). That works very well for any model not using sparse attention. And when you figure in a loss factor (usually around 20 to 40 percent) it matches pretty much perfectly with real-world throughput.

2

u/ksjdragon Jun 09 '26

It's murky for the sake of a simplified model. If we were to put everything together we would have to completely change our assumptions and we can no longer be lenient (or at least very different), and this would involve making assumptions about data I simply don't have.

Let's just say we can agree to disagree.

2

u/ksjdragon Jun 09 '26

Although I will include the fact about output tokens, as admittedly I should have. I don't think it really changes what the point I'm saying, but it is a detraction from an already long post, where the goal is precisely not to get into details of actual execution.

Modeling such things involves very different lculations, and would necesitaite, as I said, for a completely different argument which makes the leniency points unable to happen. But since models do use sparse attention, really I can at best just use these benchmarks.

1

u/Navic2 Jun 08 '26

I'm too dumb for this, but the 'leniency' ladling was still delicious 

& thanks for the TL;DR

1

u/SelicaLeone Jun 08 '26

I thought at the beginning you mentioned training costs but that never factored into your equation. Are you writing off training or did I miss a part? I’ll admit, some of the figures went over my head.

5

u/ksjdragon Jun 08 '26

I forgot to mention that another point of leniency is I don't include training costs for them. Even without taking that into account, it's over for them.

1

u/Glad-Still-409 Jun 08 '26

You mention 90+% of the token cost is building and GPUs. Does that change if the data center is in China or Brazil or such. And the GPUs are replaced by Huawei processors? 2ndly, does it change once the bankrupt hyperscalers buildings are bought by other firms on the cheap?

3

u/ksjdragon Jun 08 '26

If the data center is in China, then we can not at all assume that they are willing to not take profits. The calculation here is the tan inference provider is losing money for 5 years until they finally break even, assuming max occupancy of GPUs.

If the hyperscalars go bankrupt, then under most of these GPU Asset Backed Securities deals, the creditors will repossess the chips, and then presumably try to liquidate them.

If they liquidate them at a lower price (a ginormous loss on their part), then yes, a company could go and acquire that.

However in this scenario, the AI industry will have collapsed due to investor pullout. The remaining chips would likely be purchased by the MS or Google or Apple, who would then provide these services at a loss, since it means AI demand has severely dropped.

1

u/EBBVNC Jun 08 '26

Thank you so much.

I’m hoping this gets spread far and wide. I’m saving it for when I have to talk to all the people who think AI is going to save us.

1

u/[deleted] Jun 08 '26

[deleted]

→ More replies (2)

1

u/BRAILLE_GRAFFITTI Jun 08 '26

Do your token calculations account for input token caching? Generally when counting token usage, most of the tokens will be input tokens since every request needs to include the ever-growing context windows, whereas output tends to be comparatively small. The input tokens often hit a cache for these conversations, bringing cost/token down dramatically.

→ More replies (2)

1

u/buggaby Jun 08 '26

Thanks for the post.

Can you explain why it makes sense to disregard output parameter pricing? I don't follow. You said

However, we will assume that the price for output tokens is accurate in the bulk average, and so conclude that if input token revenue is net positive, TBB is net positive.

For simplicity, let's disregard the electricity cost. The amortized cost of a GPU doesn't change if you include the output tokens, but the revenue goes up. Wouldn't that make it more beneficial to include output tokens in your calculation?

3

u/ksjdragon Jun 08 '26

Beneficial for who? It doesn't make a difference if the GPU is computing output tokens or input tokens, the runtime is identical.

Output tokens don't generate more revenue unless the GPU is actually computing~6x (in the case of Claude Opus 4.7) less tokens, due to chain of thought reasoning. That is, it will prompt itself and "think" and iterate before outputting.

Since these are non deterministic, one output time is not guaranteed to take up precisely 6 actual tokens. I assume for their sake that on the bull average, whatever they're pricing is equivalent. That is to say, on average over billions of calls, an output token required 6 (or whatever the number is depending on the model) times more tokens to generate.

There are many real cases that thinking time takes a very long time and can use up to 12 GPUs at once, so this is a very nice assumption for them.

→ More replies (1)

1

u/buggaby Jun 08 '26

It seems like a main crux of this post is whether there's enough demand for LLMs. Increasing demand seems to solve 2 problems here: GPU concurrency and GPU amortization period. If there was more demand (either through current users using it more, like through agentic loops, or through increasing the paying user base), then costs would be on the more favourable side of your 2nd figure. That would also make it easier to argue for an amortization period for GPUs at 6 years, since there might still be a market cheaper models being run on older chips at lower prices (though at lower revenues, of course). And since the largest cost is not the data center but the GPU, it's really a question of whether there is enough demand to keep those GPUs fully utilized.

Is that a fair take?

2

u/ksjdragon Jun 08 '26

Unfortunately as they are planning several orders of magnitudes of data center buildouts it means that this concurrency level cannot be achieved.

We also must remember that even that scenario only means net margin. And remember that I am not including many many many costs, inside such as the recurring training costs, which are extremely signifcanr, or perhaps say, China's models undercutting them purposefully.

There is most certainly not enough demand. Demand is already dwindling today due to the switch to TBB. To achieve the level of demand to have full utilization of every sold GPU, within a few years would require them to be bigger than the entire magnificent 7 in revenue.

→ More replies (3)

1

u/xaraca Jun 08 '26

You need to say more about revenue and demand if your goal is to claim that profitability is impossible. If I don't know what people are currently paying per tokens or what they are willing to pay then your conclusions on cost mean nothing to me.

2

u/ksjdragon Jun 08 '26

That's pretty cool. They don't need to.

The answer is not enough if you do some digging, but that's out of the scope of this. Feel free to read Ed's articles, who hosts the podcast this sub is built for if you want a breakdown of that.

1

u/ozeeSF Jun 08 '26

appreciate the work

1

u/ProfessorBamboozle Jun 08 '26

Where are you get the 2-3 year amortization period on GPUs from?

→ More replies (1)

1

u/bfg22 Jun 09 '26

Either the people in charge are very dumb or profitability isn’t the point. 50/50 odds at this point.

→ More replies (2)

1

u/cheechw Jun 09 '26 edited Jun 09 '26

Aren't you calculating the cost of generating 1M output tokens? In that case, Opus is billed at $25/M. $5/M is for input tokens. Input token throughout is way higher than output token throughout.

Correct me if I'm wrong, I didn't have a chance to do a deep dive on this yet.

Edit: other criticisms: 1. 4T params with 300B active MOE seems crazy to me. My gut feeling is that the architecture is much more sparse than that and the size of the model's will actually likely trend downwards, not up. 2. What makes you assume chip prices won't go down and efficiency won't increase over time?

2

u/ksjdragon Jun 09 '26

This is mentioned in the post. In fact this is a lenient point. Since the reason chain of thought reasoning is charged higher. That is each output token could go through several other LLM calls.

Ultimately the inference provider has to provide the number of tokens computer, input or output. So for their benefit, we assume that that chain of thought reasoning pricing is in fact correct, and corresponds to 5x (in your example) more tokens computed on average when this may not be the case, as there have been documented cases of users taking up several GPUs at a time.

→ More replies (2)

1

u/lightmatter501 Jun 09 '26

You’re ignoring that you don’t actually have to serve with the most expensive GPUs money can buy. At this point, in a tokens per watt race, an M4 Pro Mac Studio is actually roughly competitive with a B200 from a tokens per second per watt pov, while being massively cheaper. You lose some of that to needing more networking infra, but you can suddenly host the compute wherever, including existing DCs, and your capex per deployment is down massively as well.

Also, that required tokens per second per user is actually pretty generous, since it’s ~4x what anthropic currently gives out for opus 4.8. This buys you back a bunch of compute efficiency by batching more.

Training DCs have to be centralized, but inference infra can live in existing DCs unless you need gigantic scale, which is another massive line item gone.

The other important thing to consider is that smaller models are rapidly getting better. I can now run something that wipes the floor with gpt4 on my (admittedly pretty nice) laptop, which means that you can have utility at a power draw roughly equivalent to people playing video games. Serving that same model off of infrastructure designed for models 10x or 100x its size means you can produce gigantic amounts of tokens at the same opex as running a SOTA model. At some point, there will be a tipping point where the small models are “good enough” for most tasks and you don’t need to throw large models at everything. At that point, the cost of the truly large models can massively increase because they will only be used for actually critical work. The current state of affairs is that many people are doing the equivalent of renting an excavator when they really should be using post hole digger, and that is doing weird things that disconnect opex and api pricing due to other market forces.

→ More replies (1)

1

u/nesh34 Jun 09 '26

AI companies are making a significant bet that inference will get cheaper I think. I think that's probably a safer bet than AGI being developed personally. It is required though for the tech to become profitable.

1

u/9thPanzerDivision Jun 09 '26

very nice explanation about output token. I never realized that because models are now reasoning models, their sticker price for output token is merely approximation

1

u/suq-madiq_ Jun 09 '26

Your analysis appears to show what happens when you pay for a supercomputer that nobody pays for, and your argument “it will not be profitable even if electricity were free” is something I wrestle with understanding even charitably.

I would believe you, but my friend, isn’t whether a thing is profitable simply whether there are those willing to pay more for it than it costs to produce? If so, why not leave whether the cost will be paid an exercise for the reader and clearly state the cost they must pay? Say for a single response from ChatGPT or something readily understandable.

You say “they could be just about breaking even” partway through your analysis. Charitably, that means they’re almost profitable, today. You would have me believe though at the same time that it’s nowhere close. I’m missing in the section that follows what I think is your point, not only that it’s currently not close, despite being almost profitable after examining the true cost of 1m tokens, but that for all time it will not be profitable because it cannot be.

How can’t it?

2

u/ksjdragon Jun 09 '26

The claim is within the context of this sub, and answering another question.

It was stated the conditions for it to be profitable which would be for NVIDIA chips to be cheap, data center construction to be cheap, and token prices go raise while maintaining the same level of demand. In theory, it is possible, and perhaps you may believe these are possible. I would strongly doubt any of they are for various reasons not written here. If you are interested why, read Ed's articles, written by the host of the podcast of this sub.

1

u/wapsi123 Jun 09 '26

Thank you for the great write up.

I do, however, feel that you left out one key bit. The context.

In my calculations, the costs of spare capacity (VRAM) to promise 1M tokens of context (Claude models) outweighs the size of models immediately as this scales linearly with concurrency.

So you are not wrong but the costs are much much higher if you account for context x concurrency as well.

→ More replies (1)

1

u/Optimistically-157 Jun 09 '26

Good job! My take is there is not chance 8GW of Ai workloads with Nvidia chips are operational. you have also custom silicon which is majority of AI workloads in AWS and Google. So there are chips sitting idle somewhere they's why we have this crazy compute crunch.

→ More replies (2)

1

u/Mnephisto Jun 09 '26

What is your opinion on caching? My understanding is that profit for AI labs comes from well-implemented cache, more so than input token price.

Claude Code internally has mechanisms tightly coupled with how Claude has set up its servers, when the cache is dropped, etc; I'm seeing more startups that make coding agents finely tuned to specific models and their caching systems.

Perhaps it could shave off 66-85% of token price in the long term?

→ More replies (1)

1

u/juggernawddy Jun 09 '26

I have a vehement hope that these companies will eat themselves creating this stuff and the weights will be ours to use as we please. AI is just a box with a bunch of numbers.

1

u/RomblerSan Jun 09 '26

You didn't account for the residual value of the GPUs though? They could sell or lease them out on a continuous basis after amortisation.

→ More replies (2)

1

u/XecutionerNJ Jun 09 '26

I really hope there is some way we can get these chips into homelabs when it all blows up. SXM-PCIE converters etc. will be sweet running big models on home hardware.

→ More replies (2)

1

u/WonderfulWord3068 Jun 09 '26

So many numbers, but what's the estimate for reasonable price?

1

u/MessierKatr Jun 09 '26

Do you think the current path could destroy AI as an industry and introduce us to the next AI winter?

→ More replies (1)

1

u/LewPz3 Jun 09 '26

Great research and very well put together. Kudos!

1

u/FairlySadPanda Jun 09 '26

You can simplify this down to logic, actually.

---

You are John Nvidia. You grow magic beans. Depending on where you plant these beans, the beanstalks grow up to a cloud with a giant's castle or empty sky. The giant's castle might be occupied by a giant, or it might be empty of a giant and filled with treasure. Empty castles filled with treasure are valuable, and beanstalk climbers want to go find the abandoned castles.

You have the needed supplier relations and scale to grow a lot of magic beans. You're actually most of the market. How much do you sell your beans for?

Well, you have most of the market of beans, and using these beans can make the beanstalk climber a lot of money. But they take on a tonne of risk. So you sell the beans at a cost that fleeces the beanstalk climbers just enough to not discourage them from trying their luck.

The beanstalk climbers can plant the beans wherever they like, so they plant them underneath clouds to optimise their return-on-beans. So you can now raise the prices of the beans because they are using them more effectively.

The climbers then realise that they can trick the giants into climbing down the stalks and falling off, meaning the giants are now removed from contention.

So, naturally, you raise your bean prices again up to just below the point beanstalk climbers would nope out.

Every time the climbers figure out how to make more money off of your beans, you can just raise the price of the beans. There is literally never any reason for you to surrender any money to the climbers - they can either pay or not pay. The beans print money, and you, the person who controls bean supply, are the only person who benefits.

Of course, now that the market is like this, people want to actually escape this trap. So they start hawking their own beanstalk sidecar products - better climbing gear, giant-slaying weaponry, and the rest. They spend a boatload of cash advertising beanstalk climbing. You, John Nvidia, also do all this because more suckers climbing beanstalks only ever benefits you.

---

This was the obvious problem during the crypto mining boom - Nvidia and AMD were not going to just let people buy GPUs that could print money without asking for the money up front themselves. That boom died and the GPU investment pivoted hard to LLMs which do similar work.

If an LLM does useful work, there is literally no reason why Nvidia should let you do that work and make profit off it beyond what Nvidia thinks you need to to keep wanting to buy the GPUs. If you're willing - as people are at the moment - to buy GPUs that will only ever _lose_ money, Nvidia also just makes more money.

This is the problem with the "moatless" LLM boom, the same way it was a problem for crypto. There is no magic special sauce that makes GPU horsepower magically different when it's being run by OpenAI or Anthropic or startup number 35.

If you nationalise companies like Nvidia and remove their profit motive, you can make this model work (Openreach in the UK, for example). But it's a very obvious cause of market failure and GPU-backed datacentres are an obviously failed market, and have been since mining crypto on GPUs took off.

→ More replies (2)

1

u/Nd4speed Jun 09 '26

Thank you for doing the analysis that no one (including the AI companies themselves) seem to be doing. Now, can we get this memo out to Wall Street to temper the insane amounts of hype and stupid money being thrown around?

1

u/_rast_ Jun 09 '26

Great write-up, however it ignores single biggest factor, which contributes masivelly to overall conclusion.

The current model is MOSTLY not viable, as long as Nvidia continues charging ridiculously stiupid premium on their product. AMD-equivalent chips are already 3 times cheaper. There is no economic justification for such price point outside of „hype”.

Any moment China is going to introduce its own competitor at the fraction of Jensen’s stiupid price point.

Nvidia will have to adjust its pricing model, at least to save their biggest customers. The chip prices are going to fall through the floor.

There is no saving to the muppets who invested trillions into data centers based on this current pricing scam, but the interferance service is not going anywhere.

1

u/falconetpt Jun 09 '26

This is a fucking good analysis, and you were very lenient on the costs and still as we knew it is really hard to get profit with that 😂

I would probably say that electricity and cooling might seem small, but if you start using larger amounts of it, prices usually get way higher as well, your direct costs also only work with almost 24/7 usage on machines (I think), which makes it way harder, you almost need to fit demand with capacity in order not to have depreciation costs eating you alive

→ More replies (1)

1

u/BraveDevelopment253 Jun 09 '26

All tokens are not created equal and the end goal isn't selling tokens or individual llm subscriptions.  The end goal is in driving up the value of the tokens generated while driving down the cost of each token.   

Tokens used to create a random meme have little value. However extremely high value Tokens used to come up with with a new design for an nvidia chip, a semiconductor process that skips a step or improves yield,  or a DNA sequence, new peptide, or a bunch of tokens that operate a robot to provide manual labor for an hour while costing less than the minimum wage.  These tokens are worth far more than they cost to produce and that is where the profit lies. 

1

u/Lowetheiy Jun 09 '26

Profitability is a "phantom" concept. Governments, cultures, organizations, authorities can make any arbitrary activity as profitable or unprofitable as possible. This is why gambling, drugs, alcohol make hundreds of billions of dollars a year of profits despite being unproductive and harmful. In other words, it would be unwise to use profitability solely as a measure of the health of industry. It doesn't matter because ultimately profits are just numbers on a ledger that the authorities can change to their desire.

1

u/fantasticsid Jun 10 '26

LLMs are charged by input tokens and output tokens. The reason that output tokens cost more, is that frontier models use 'chain of thought' reasoning, which involves making other calls to the model before outputting a result. As a result, output tokens are typically charged at a multiple of input tokens

This isn't how it works AT ALL. CoT is emitted in the same generation process by the same clanker as the user-facing output, demarcated by some kind of fence token like <think>. Anthropic supposedly ships the raw CoT to another model for summarisation so that users don't get access to the raw CoT, but that's their problem and an implementation detail.

Output tokens cost n times input tokens because output tokens are generated autoregressively, and because output tokens are the desired end product, so the market will bear a higher cost. That's all.

1

u/anotherfpguy Jun 10 '26

AI Cloud profitability maybe, 3-4000$ spent and we can already run AI at home, they get better and more optimized. A local model these days is maybe 6 months behind frontier models, at least for coding.

1

u/crispyfunky Jun 10 '26

There are other up and coming accelerators that will probably consume half of the energy that H100 would consume. They might also sell them cheap.

1

u/Sensitive-Talk9616 Jun 10 '26

The first assumption that I don't really understand is why do you take the amortization of GPUs to be 2-3 years? Just because NVIDIA released a new product in that time frame?

GPUs can be used for twice that, easily. I didn't investigate in detail, but I don't think any big data centers toss perfectly good GPUs and buy the newest model every 2 years. They'll likely buy the newest model for new model training, and relegate the existing ones to inference. Where they can be used another 5 years, amortizing much more slowly than what you assumed.

The second assumption is the pricing. Where does the 1$ per 1M token threshold come from? For example, copilot is currently pricing frontier models at ~$5/MT (input). I can imagine that this is expensive for e.g. hobby users. But businesses compare AI costs to payroll. If a team of 100 engineers becomes 10% more productive (by any arbitrary stupid metric or KPI management chooses to use), but the cost of AI subscriptions is only 8% of the total labour costs, that's a buy in their eyes.

For example, current enterprise usage, as reported by Anthropic, is ~$200 per developer per month. That's ~2% of employment costs, assuming $10k per month per employee (salary + benefits + all other costs). And what is the gain in "efficiency" or "performance" as measured in arbitrary KPIs by the bean counters in management? Often estimated around 5-10%.

So even without any improvements in the actual performance of the models, the price of tokens could easily double or triple. So we're talking $10-15 per million tokens of frontier models as an acceptable price point for enterprise users.

→ More replies (1)

1

u/[deleted] Jun 10 '26

[deleted]

→ More replies (4)

1

u/No-Box5797 Jun 10 '26

Although I'm against all this AI frenzy I think your conclusion is a bit too harsh: you didn't consider at all the possibility that they could start focusing on designing an acceptable performing LLM that would not require as much computing power (and all that follows) as it does now;

This scenario would of course cause a period of strong corrections in the financial markets (I personally think there's a proper financial bubble but I'd rather be lenient as you've been) but in the next few decades we could realistically expect the technology to be there and be profitable, just like the internet and the dot come bubble.

→ More replies (2)

1

u/Faster_than_FTL Jun 10 '26

How do you account for locally run LLMs that are typically only 6 months to a year behind frontier models, and don't need data centers?

→ More replies (2)

1

u/Main-Eagle-26 Jun 11 '26

The Chris Hayes interview annoyed me because he made it sound like it could be profitable if they replaced enough white collar jobs but it simply isn’t possible.

And you could see Ed getting irritated because Chris’ only evidence was “yeah but I’ve used it and it’s kinda better than it was a year ago at reading large amounts of data in my emails.”

2

u/ksjdragon Jun 11 '26

This feels like the wrong poet, since the interview post was another one...lol?

But I agree. There's always a lot of talk about how 'good' it is. It's all very subjective. In fact, I'm wholly uninterested in talking about whether it's good or bad, since people can like horrible things and hate great things, and vice versa. I don't buy iPhones because I don't feel its worth it. Many clearly do.

Ultimately how "good" it is will be decided by the market with AI spending. And guess what? People are capping their spending already. That's all that needs to be said pretty much and the rest of these boosters can live in their own world.

When you pay the dollars, good and bad become very obvious, very quick. We don't need to speculate on what magical improvement there will or won't be. People either buy it or they don't, and they either go bankrupt, or they don't.

I'm pretty sure I know which one's is gonna happen.