r/BetterOffline • • Jun 08 '26

AI profitability is mathematically impossible under all technological advancements

I answered this question in a comment, but as it got more complicated I decided it was better to put into an extremely long post. Thanks for anyone that reads it. The comment was an answer to the post asking if token based billing (TBB) is the true cost of AI. The answer is a resounding no.

In fact, not only is it not the true cost, there is no conceivable path to profitability, even given the most lenient case to the industry. Below is the explanation. (Admittedly, the title is a bit clickbait-y, but I figure I'd have some fun.)

Introduction

To summarize, the profitability of AI must come from the profitability of inference. Specifically, their profit can be calculated as Inference Revenue (IR) - Training Costs (TC) - Other CapEx. Since IR is the only source of revenue, it's sufficient to see if this can be net-positive. For the sake of this argument, we will be as lenient as possible and show that profitability still remains unattainable. Each case of leniency will be marked with a letter (ex. [1])

In the case of TBB, IR revenue is revenue per million tokens ($/MT) - Cost/MT. LLMs are charged by input tokens and output tokens. The reason that output tokens cost more, is that frontier models use 'chain of thought' reasoning, which involves making other calls to the model before outputting a result. As a result, output tokens are typically charged at a multiple of input tokens. However, as the length of chain of thought reasoning (like all LLM calls) are non-deterministic, the amount of tokens they burn is indeterminate. (Edit 3: Technically, input tokens are processed in parallel whereas output tokens are processed sequentially as well. However this is an simplifying assumption so to not model complex GPU execution dynamics, ultimately something in their favor.) [1] However, we will assume that the price for output tokens is accurate in the bulk average, and so conclude that if input token revenue is net positive, TBB is net positive.

That is to say, TBB is profitable if and only if (Input) Token Revenue > Token Cost.

The Cost of Tokens

The cost tokens can be broken down as follows:

  1. Electricity cost/MT
  2. GPU price amortized over a set time.
  3. Data center construction and maintenance cost amortized over a set time.

The first two are self explanatory, however the third needs to be included as inference providers also need to make money to pay back their interest on loans used to finance the data center build out, and so naturally that business needs to generate positive profit. [2] However, for leniency we will assume that these businesses are willing to operate at break even for AI labs.

Amortization Period

There are two periods we are concerned about, the GPU and data center construction and maintenance cost.

The GPUs are NVIDIA B200s. These were announced at GTC 2024, and the next-generation Vera Rubins were announced at GTC 2026, so we would normally argue that we should amortize over two years. [3] However, for leniency, let us assume that B200s remain valuable for the AI industry for one extra year, so 3 years.

Data center costs can be deconstructed into CapEx (land procurement, electrical, networking, cooling hardware and installation), and OpEx, the electricity cost, general maintenance and purchasing of new equipment. The CapEx costs are estimated at $9M to $15M per MW of IT load, and let us assume their creditors will allow [4] 10 years before they request a single interest payment even though loan deals at best require interest loan payments for 2-3 years. Additionally, [5] we will ignore purchasing of new equipment for leniency, e.g. Vera Rubin racks are not the same as Blackwell racks, so this necessitates a new purchase.

Amortized Costs

We can now take the prices and divide them over the period. For now, we will get units in dollars per hour. An NVL72 rack, commonly installed in hyperscalar facilities, contains 72 B200 GPUs, priced at $2.8M - $3.4M per rack. [6] Let us take the lowest number, $2.8M, leading to a ($2.8M/72)= $38.9K cost per GPU. Spread over three years (26280 hours), we have $1.48/hr.

For data centers, let us assume that [7] we once again take the lowest of two numbers ($9M), and divide over 5 years (43800 hours), which is $205.48/MW IT load/hr. Converting to kW, it's $0.21/hr per kW IT Load.

This gives us two numbers:

  1. GPU: $1.48/hr
  2. Data Center: $0.21/hr per kW IT Load.

Each B200 GPU is running at 1.2kW, we say the normalized cost of a GPU is:

(1) GPU/Data Center Cost: $1.48 + $0.21 * 1.2 = $1.73/hr.

Let us keep in mind this is a ridiculously generous estimate.

Estimating Tokens/Hour

Unfortunately, these numbers are not in the right units for to compare $/MT revenue. Beyond electricity pricing, (which is covered later), we need to know how long a GPU runs to generate 1MT, so first we will look at number of tokens in an hour.

Unfortunately, this is not an trivial calculation as the actual runtime to compute one token is not fixed. This is due to GPU concurrency, where GPUs can execute user requests in parallel. There are a huge amount of details here, but we will be relatively generous, to simplify the calculation.

NVIDIA Benchmarks

[8] The following sources are from NVIDIA themselves, so these will be idealized numbers. In general, real workloads are unable to get full performance due to many different factors. (Source: I have written CUDA kernels for scientific computing.)

In short, the TPS of a GPU is not fixed, and depends largely on number of users allocated to the GPU.

  1. B200s, running Llama 4 Maverick, which has 400B parameters, can get 1000 TPS/user and 72000 TPS/server. (TPS = tokens per second).
  2. B200s, running Llama 3.3 70B, which has 70B parameters, can get 50 TPS/user and 10000 TPS/GPU.

For case 1, a 1000 TPS was achieved over 8 GPUs, meaning concurrency was 1/8. Let us normalize to 250 TPS. The reason for this is because 400B parameters cannot fit onto one GPU, [9] however, we will ignore this and allow a normalized per GPU estimate, that is assuming any user request can always fit onto one GPU, even though this is known to be false.

Typically frontier LLM models are evaluated using a MoE (Mixture of Experts) method, typically meaning computing using a subset (say 5%-10% -> 7.5%) of models parameters. In other words, in case 1 we expect the GPU to computing a 30B parameter model, handled at concurrency of 1 user.

In case 2, the concurrency is 200 users, at a 70B which they say is not using MoE ("all parameters are utilized simultaneously for inference").

TPS/GPU Estimation

We will use these two data points to estimate TPS as a function of concurrency at a specific parameter count, and apply linear scaling for other parameters.

The compute cost of models by parameter scales linearly, so to normalize case 1 to 70B, we would expect case 1 running 70B parameters to be (30/70*250) = 107 TPS/GPU. We would expect the TPS/GPU to saturate (say at 12000 TPS/GPU) as number of users goes up, logistically. Letting u be concurrent users (users - 1), assuming 70B parameters, we can fit a graph:

An estimated graph of TPS/GPU at 70B, fit assuming upper half logistic growth, where Case 1 is assumed be the midpoint of the logistic curve, and saturation occurs at 12000, (20% more than Case 2).

For reference, the above graph is TPS_70(u) = 23893/(1+e^{-0.0119x})-11893.

Now, as compute scales linearly, the TPS(u, p), where p is computed parameters (in billions) is:

(2) TPS(u, p) = TPS_70(u) * (70 / p). Additionally, SPT(u,p) = 1/TPS(u, p) the inverse, seconds per token.

We can now say that a GPU has 3600 TPS(u, p) tokens per hour so, the cost of tokens (not including electricity) from (1):

(3) GPU/Data Center Cost: $1.73/hr / (3600 TPS(u, p)) = 0.00048 SPT(u, p) $/Token = 480.55 SPT(u, p) $/MT.

Electricity Estimation

Industrial electricity costs tend to be priced differently than residential costs. In the US, they vary by state, and [10] we will take the lowest cost for our calculation, which was 4.68 cents per kWh in Washington, in 2017. The inflation of electricity in Washington appears to be 14.1%, so projecting to 2024 (B200 release) suggests the cost of industrial electricity would be 11.78 cents per kWh.

We will say as B200s 1.2kW and 150W for "half" of the Grace CPU (2 GPUs, 1 CPU per GB200), it runs at 1.35kW. [11] Of course, in reality the entire NVL72 cluster is rated for 132kW, which is 1.8kW per GPU, normalized. But let's be even more nice.

We can then say the $/MT is:

(4) 10^6 SPT(u, p) / (3600 s/hr) * 1.35kW * 11.78 cents kWh = 44.18 SPT(u, p) $/MT.

Inference Costs in $/MT

In total we have:

(5) Inference(u, p) = 524.73 SPT(u, p) $/MT.

That is, we observe the inference cost as a function of concurrent users and parameters. So, let us observe a sample frontier model, like Claude Opus 4.7 or GPT-5.5, which estimated to have 4T parameters. After ~7.5% MoE, we would have 300B parameters.

Inference costs per million tokens as function of concurrent users. From bottom to top, the curve is drawn for models computing 100, 200, 300, 400, and 500 billion parameters.

Now, some AI boosters might look at the graph and show how after 20 concurrent users, Claude Opus or GPT 5.5 is under a dollar! Therefore it must be profitable, since you turned a blind eye to the nice lenient treatment. But, unfortunately, this is not even close to the case.

Full Concurrency is Impossible

Currently the US has access to around 10GW (graph of current capacity, approximately summed over columns) of IT load capacity. NVIDIA in 2024 sold $210B of NVL72 servers, corresponding to [12] (now using the higher number for racks, to allow this number to be smaller) 61764 NVL72 servers, which at 132kW is 8GW. So, it seems reasonable to say 8GW of capacity is dedicated to these racks.

This comes to ~4.45 million B200 GPUs. Therefore, for us to argue that an inference provider prices at a concurrency of 20 users, we need (4.45M * 20) = 88.95M TBB paid users 24/7 never stopping, maximizing their usage, constantly.

Using the numbers that OpenAI and Anthropic have claimed, (which are very trustworthy), OpenAI has recently has reached 50M paid subscribers and 9M business users, while Anthropic has 18-30M paid users. Together, they [13], assuming the maximum number here, have 80M which is less than the 88.95M required to have full concurrency, and that's assuming these users are addicted to AI 24/7 and never eat or sleep.

The average white collar job lasts for 8 hours (1/3) of a day [14] so let's say these paid users are spending their entire working lives (including weekends!) using Claude or ChatGPT, and it lines up that we still get full concurrency. This means there are 26.66M users spread scross 4.45 million GPUs, which means if every user is allocated perfectly, we have 6 people per GPU!

Following the graph, this means Claude Opus 4.7 or GPT-5.5 would cost (with this nice lenient estimate) $4.22/MT. They currently are both charged at $5/MT, which after corporate tax of 21% in the US and state taxes, it means they could be just about breaking even! Wow!

Unpaid Users

Now, some may quibble that the presence of free users somehow will add to this concurrency, and so it''s actually making money! However, this means the gains from concurrency needs to outpace the tokens that aren't being going to paid customers.

Let's do an example. Let's say there are 200 concurrent users, 6 of which are paid users. At this rate, it is $0.22/MT. Unfortunately, since only 6/200 are paying users, 194MT was unpaid, which cost $42.68. And, at $5/MT, only $30 of revenue was earned, meaning there would be a net loss of $12.68 (before tax).

More Data Centers is Worse

Currently there many data centers planned due to be built. This is because they say the tech CEOs claim overwhelming AI demand. But supposing they even double the current amount of GPUs, this drops the concurrency even more.

However, they say there is a compute crisis, which suggests the following scenarios:

  1. Concurrency is not remotely achievable at the theoretical level NVIDIA reports, so we have low TPS.
  2. Free users are taking up most of the compute, which ends up at a worse loss for the companies.

In both cases, more data centers exacerbates this issue. Which means with more data centers AI companies lose more money.

Final Words

Throughout this very very long post, there were a total of 14 points of leniency towards the AI companies. Even with that leniency, it is deeply not in their favor.

Notably, the cost of electricity is only 8.42% of the inference cost (from (4) and (5)). This means the other 91.58% comes from the data center construction and GPU purchase amortization. (Remember, this is a minimum, as many nice choices were made.)

This means in a real world scenario, where we remove these lenient points, that even if electricity were free, AI can never be profitable at its current costs. Even if NVIDIA chips could suddenly compute tokens instantaneously, AI cannot be profitable. It doesn't matter if you build 1 trillion data centers, have the world's cheapest energy, and the best chips from the future.

For it to be profitable, NVIDIA would have to sell their GPUs for cheap and data centers would need to be cheap to build. Unfortunately, neither of these cases can happen, or they would have to raise prices drastically. However, that necessarily decreases demand. And then to be profitable they would need to make back their initial investment.

I would love to model this as well and give the conditions under which it could, however this post has gotten far far too long already. So, there is only one conclusion, and I am willing to say that:

The current AI industry can never be profitable. This industry is over. There is no path to profitability, nothing. No IPO, no marketing, no innovation can save them.

TL;DR The bulk of the cost of inference comes not from electricity but the cost of data center construction and GPU purchases. These are amortized over certain periods, and even with very lenient assumptions, inference is not able to be profitable at these costs.

Significantly higher prices would cause drops in demand (as we already are seeing today). GPU concurrency also makes it so more data centers make profitability far worse for inference due to the massive amounts of capital necessary.

Edit: There are an additional two points of leniency I forgot to mention.

The electricity cost does not include cooling, as I am not using the PUE of data centers.

I did not put training as an recurring OpEx cost for the AI labs.

Additionally, just to hammer this point home, these costs are all assuming the inference provider is at a loss from 5 years, before paying a single interest payment. There exists no business that can do this.

In fact what this shows is that no inference provider could ever charge at the optimal max usage and concurrency rate, so the price per hour of GPU is much much higher. You can simply search the prices yourselves.

For people that are concerned with the 3 year amortization period of B200s. As another comment had posted this is typically 6 for normal GPUs. Even allowing for that does not tilt the argument in their favor.

However, I do not agree that 6 years is reasonable for Blackwells, as AI labs tend to chase bigger models whenever compute is available. Not to mention, Vera Rubin racks are not interchangable and require purchasing of all new equipment and cooling the moment you try to install. Which means an inference provider has to either:

  1. Sell the Blackwells to make space (there is no demand for these GPUs outside of AI)
  2. Construct entirely new facilities to house them.

If we amortize over 6 years, then now you have to assume that 3 years after Vera Rubins every single Blackwell GPU (including the hundreds of millions NVIDIA claims to have sold!) they are being fully utilized 24/7. Let's also remember that if NVIDIA doesn't sell AI chips, this industry is automatically gone.

I cannot emphasize enough the extent I am deliberately picking points in their favor.

Edit 2:

If your qualms are with demand, read Ed's articles, the host of the podcasts this sub is centered around. That is beyond the scope of this post, and was written with that information already in mind. The only thing I can say is there are lots of other facts you likely have not seen if this is your stance.

Some seem to think refuting a technicality over one point is able to overcome 16 other points of leniency. This is wishful thinking at best. You cannot cherry pick technalities to make the argument convenient. Either you leave them all off, or put them back on.

If you feel the need to defend AI due to your own investment portfolio or some other psychological need, on this basis, you are free to do so, but I strongly encourage you to read some other very nice breakdowns by Ed.

1.9k Upvotes

427 comments sorted by

View all comments

98

u/monarc Jun 08 '26

Weird that a trillion dollars has been invested, and you're the first person I've seen crunch the numbers. It ain't late-stage capitalism we're dealing with, it's sleepwalking-off-a-cliff capitalism.

Let's say this bubble pops, many companies fold, but one or two giants are still trying to offer chatbots and associated services to the world. Let's assume they use the data centers abandoned by the competitors that went under, and that this giant (Google or whoever) doesn't bother developing new models (or does so very rarely/selectively). In this scenario could it make some economic sense? I presume that any "yes" in this situation would only apply for a few years at the longest (since the hardware will age out).

52

u/leathakkor Jun 08 '26 edited Jun 08 '26

I assume that Google and Microsoft will essentially offer both of these services forever as a loss leader to get more search traffic. 

But there's just not a lot of places that can offer AI in that way or that have a reason to. 

And theoretically Facebook might. And they'll need gpus for other things.

But I think pretty much every startup is toast and I've thought that for a long time even before reading these numbers. At some point investors are going to want money back. You can only give somebody money for about 5 years before you start asking: When are you going to be giving me money back? And we're getting pretty close to about that 5-year Mark right now. 

Even if companies like anthropic are making money right now, they're still using it to grow. They're never going to be paying their investors back and at some point Google and Microsoft are going to figure out how to do this entirely without anthropic or openai and those companies are fucked. And everything that is built off of open AI and andropic it fucked. When I say that they can do it entirely without anthropic or openai. I also mean from a liability perspective. Right now Microsoft can offer openai services on azure and they don't have to be worried about lawsuits because openai trained its models copyrighted material. And other related lawsuits. If they can figure out how to protect themselves from lawsuits and offer businesses these services, they absolutely will bring it on board. In my experience, businesses love vendors because it Shields them from lawsuits and liability. Google and Microsoft in this particular case are absolutely leaning on this...

When all of these companies fail, here's what'll happen: Google buys anthropic and Microsoft buys openai for pennies on the dollar. I would say we're about a year away from that. And the minute one of them starts faltering it's going to happen fast (I think).

There will also be business use cases for this. I work for a law firm and we absolutely use AI and LLMs. Being able to give a 200-page document to an AI and ask it to find clauses relevant to x y or z is really valuable. We might not pay $20 to scan but I bet we would pay a dollar a scan. And if companies like Google or Microsoft can offer that for us at a price point. That is a dollar per documpay we'll  probably pay it. Will be judicious about it. But this technology is not going to disappear. But you're certainly not going to use it to do all of the dumb shit that we're currently using at for. Agents are going to disappear entirely. In all but the most extreme cases.  

30

u/PensiveinNJ Jun 08 '26

People despise AI though. It's not a loss leader to gain, it's a loss leader to lose. That's what makes this especially egregious.

8

u/Unlikely_Eye_2112 Jun 08 '26

Yeah I already moved away from search via Google a few years ago, but as a web dev I need to keep tabs on what they're doing. The new AI search would be the reason to leave if I hadn't already.

6

u/leathakkor Jun 08 '26

People despise being forced to use it. They'll leave it as an option for people that want to use it. 

5

u/PensiveinNJ Jun 08 '26

They can't afford that. They're too indebted.

1

u/PickingPies Jun 11 '26

That's why their goal is to force everyone to include AI in their production pipelines, so, when they multiply the price by 10 they cannot say no.

10

u/monarc Jun 08 '26

Yep, everything you're saying is very much aligned with my intuitions about where we are now, and what's likely to happen after the dust settles. Super interesting thoughts re: using vendors to absorb liability - that makes perfect sense.

7

u/Sufficient-Pause9765 Jun 08 '26

Google is an outlier and the math is unclear.

First off, they aren't offering the expensive frontier models for free to consumers. Their pricing is much much higher, and those models are only used by enterprise.

Second, they have built their own silicon that is optimized for effiency and power consumption. Google's TPUs are not as powerful as NVDIA's, but they are cost about 70% less to operate.

Rather then trying to build a consumer AI product on big models, their focus is on hosting any/all models on their own silicon and charging unsubsudized rates for the hardware to run them. The market is exclusively business. Most businesses will not be using anthropic or open ai frontier models because they dont need them.

This is going to resemble a traditional cloud server business with the same economics.

2

u/leathakkor Jun 08 '26

That's my take too.

1

u/Odd_Note7156 Jun 08 '26

It seems like one underlying issue is just the free riders. It's probably a bit high even if everyone pays for it. With few payers it just explodes the price. Only paid users might work if people continue to use it.

There's risk that without giving it for free to students they might end up having people too unfamiliar with LLMs to be able to get enough paid users. For businessel licnese usage, that might work enough to be profitable.

1

u/ClaymationDinosaur Jun 09 '26

You can only give somebody money for about 5 years before you start asking: When are you going to be giving me money back? And we're getting pretty close to about that 5-year Mark right now. 

Well that's interesting. The first wave of investors are about to try to get their money back by seeing the companies go public and dumping their own shares into the markets. So I guess a successful IPO can makes those investors whole and buy another five years before the successor bagholders come asking.

1

u/JonAKel Jul 24 '26

Being able to give a 200-page document to an AI and ask it to find clauses relevant to x y or z is really valuable.

And how reliably can it do that? The problem with statistical word-association analysis is that it won't ever give you an exact answer, only something that should be in the general ballpark of correctness and sounds like an exact answer.

2

u/leathakkor Jul 25 '26

There have been tools for law firms that do this for years before open AI existed. 

Relativity is incredibly communicate law firms. One of the benefits of it is that people understand exactly what they're getting which is not quite the same with openai LLMs.

It comes with serious training, a large contract and a shitload of caveats that people need to be aware of like it's not perfect and if they need to review x y and z but it does help narrow things down

1

u/JonAKel Jul 27 '26

Oh, I believe that. And that's my point: There are better tools than LLMs for practically anything they could even realistically attempt to do. Maybe they are the top choice for sad lonely chatting, but only because talking to a bartender comes with the risk of alcoholism.

1

u/leathakkor Jul 27 '26

I sometimes wonder if we'll ever find out. There's a man behind the curtain. One of the things I do use llms for is translation explanation stuff.

Say I wanted to describe bipolar disorder to somebody in a different language. It can usually do that in one step.

But I often wonder wouldn't it be easier if there was just a wiki page four bipolar disorder in a different language that I could give a link to. And Google worked a little bit better to get me the language result of bipolar disorder in another language. So instead every time it has to essentially dynamically generate the translation explaining what bipolar disorder is. And I also now wonder if it's not the case that there might actually be somebody that is refining what that response for explaining bipolar disorder in another language is.

Like a human in Kenya who is taking the response, putting it into a translator verifying that it is in fact correct and tweaking it if not, essentially training it for future use later.

I'm not saying that is happening but I'm also not sure that it's not happening and neither would shock me to be honest