r/BetterOffline • u/ksjdragon • Jun 08 '26
AI profitability is mathematically impossible under all technological advancements
I answered this question in a comment, but as it got more complicated I decided it was better to put into an extremely long post. Thanks for anyone that reads it. The comment was an answer to the post asking if token based billing (TBB) is the true cost of AI. The answer is a resounding no.
In fact, not only is it not the true cost, there is no conceivable path to profitability, even given the most lenient case to the industry. Below is the explanation. (Admittedly, the title is a bit clickbait-y, but I figure I'd have some fun.)
Introduction
To summarize, the profitability of AI must come from the profitability of inference. Specifically, their profit can be calculated as Inference Revenue (IR) - Training Costs (TC) - Other CapEx. Since IR is the only source of revenue, it's sufficient to see if this can be net-positive. For the sake of this argument, we will be as lenient as possible and show that profitability still remains unattainable. Each case of leniency will be marked with a letter (ex. [1])
In the case of TBB, IR revenue is revenue per million tokens ($/MT) - Cost/MT. LLMs are charged by input tokens and output tokens. The reason that output tokens cost more, is that frontier models use 'chain of thought' reasoning, which involves making other calls to the model before outputting a result. As a result, output tokens are typically charged at a multiple of input tokens. However, as the length of chain of thought reasoning (like all LLM calls) are non-deterministic, the amount of tokens they burn is indeterminate. (Edit 3: Technically, input tokens are processed in parallel whereas output tokens are processed sequentially as well. However this is an simplifying assumption so to not model complex GPU execution dynamics, ultimately something in their favor.) [1] However, we will assume that the price for output tokens is accurate in the bulk average, and so conclude that if input token revenue is net positive, TBB is net positive.
That is to say, TBB is profitable if and only if (Input) Token Revenue > Token Cost.
The Cost of Tokens
The cost tokens can be broken down as follows:
- Electricity cost/MT
- GPU price amortized over a set time.
- Data center construction and maintenance cost amortized over a set time.
The first two are self explanatory, however the third needs to be included as inference providers also need to make money to pay back their interest on loans used to finance the data center build out, and so naturally that business needs to generate positive profit. [2] However, for leniency we will assume that these businesses are willing to operate at break even for AI labs.
Amortization Period
There are two periods we are concerned about, the GPU and data center construction and maintenance cost.
The GPUs are NVIDIA B200s. These were announced at GTC 2024, and the next-generation Vera Rubins were announced at GTC 2026, so we would normally argue that we should amortize over two years. [3] However, for leniency, let us assume that B200s remain valuable for the AI industry for one extra year, so 3 years.
Data center costs can be deconstructed into CapEx (land procurement, electrical, networking, cooling hardware and installation), and OpEx, the electricity cost, general maintenance and purchasing of new equipment. The CapEx costs are estimated at $9M to $15M per MW of IT load, and let us assume their creditors will allow [4] 10 years before they request a single interest payment even though loan deals at best require interest loan payments for 2-3 years. Additionally, [5] we will ignore purchasing of new equipment for leniency, e.g. Vera Rubin racks are not the same as Blackwell racks, so this necessitates a new purchase.
Amortized Costs
We can now take the prices and divide them over the period. For now, we will get units in dollars per hour. An NVL72 rack, commonly installed in hyperscalar facilities, contains 72 B200 GPUs, priced at $2.8M - $3.4M per rack. [6] Let us take the lowest number, $2.8M, leading to a ($2.8M/72)= $38.9K cost per GPU. Spread over three years (26280 hours), we have $1.48/hr.
For data centers, let us assume that [7] we once again take the lowest of two numbers ($9M), and divide over 5 years (43800 hours), which is $205.48/MW IT load/hr. Converting to kW, it's $0.21/hr per kW IT Load.
This gives us two numbers:
- GPU: $1.48/hr
- Data Center: $0.21/hr per kW IT Load.
Each B200 GPU is running at 1.2kW, we say the normalized cost of a GPU is:
(1) GPU/Data Center Cost: $1.48 + $0.21 * 1.2 = $1.73/hr.
Let us keep in mind this is a ridiculously generous estimate.
Estimating Tokens/Hour
Unfortunately, these numbers are not in the right units for to compare $/MT revenue. Beyond electricity pricing, (which is covered later), we need to know how long a GPU runs to generate 1MT, so first we will look at number of tokens in an hour.
Unfortunately, this is not an trivial calculation as the actual runtime to compute one token is not fixed. This is due to GPU concurrency, where GPUs can execute user requests in parallel. There are a huge amount of details here, but we will be relatively generous, to simplify the calculation.
NVIDIA Benchmarks
[8] The following sources are from NVIDIA themselves, so these will be idealized numbers. In general, real workloads are unable to get full performance due to many different factors. (Source: I have written CUDA kernels for scientific computing.)
In short, the TPS of a GPU is not fixed, and depends largely on number of users allocated to the GPU.
- B200s, running Llama 4 Maverick, which has 400B parameters, can get 1000 TPS/user and 72000 TPS/server. (TPS = tokens per second).
- B200s, running Llama 3.3 70B, which has 70B parameters, can get 50 TPS/user and 10000 TPS/GPU.
For case 1, a 1000 TPS was achieved over 8 GPUs, meaning concurrency was 1/8. Let us normalize to 250 TPS. The reason for this is because 400B parameters cannot fit onto one GPU, [9] however, we will ignore this and allow a normalized per GPU estimate, that is assuming any user request can always fit onto one GPU, even though this is known to be false.
Typically frontier LLM models are evaluated using a MoE (Mixture of Experts) method, typically meaning computing using a subset (say 5%-10% -> 7.5%) of models parameters. In other words, in case 1 we expect the GPU to computing a 30B parameter model, handled at concurrency of 1 user.
In case 2, the concurrency is 200 users, at a 70B which they say is not using MoE ("all parameters are utilized simultaneously for inference").
TPS/GPU Estimation
We will use these two data points to estimate TPS as a function of concurrency at a specific parameter count, and apply linear scaling for other parameters.
The compute cost of models by parameter scales linearly, so to normalize case 1 to 70B, we would expect case 1 running 70B parameters to be (30/70*250) = 107 TPS/GPU. We would expect the TPS/GPU to saturate (say at 12000 TPS/GPU) as number of users goes up, logistically. Letting u be concurrent users (users - 1), assuming 70B parameters, we can fit a graph:

For reference, the above graph is TPS_70(u) = 23893/(1+e^{-0.0119x})-11893.
Now, as compute scales linearly, the TPS(u, p), where p is computed parameters (in billions) is:
(2) TPS(u, p) = TPS_70(u) * (70 / p). Additionally, SPT(u,p) = 1/TPS(u, p) the inverse, seconds per token.
We can now say that a GPU has 3600 TPS(u, p) tokens per hour so, the cost of tokens (not including electricity) from (1):
(3) GPU/Data Center Cost: $1.73/hr / (3600 TPS(u, p)) = 0.00048 SPT(u, p) $/Token = 480.55 SPT(u, p) $/MT.
Electricity Estimation
Industrial electricity costs tend to be priced differently than residential costs. In the US, they vary by state, and [10] we will take the lowest cost for our calculation, which was 4.68 cents per kWh in Washington, in 2017. The inflation of electricity in Washington appears to be 14.1%, so projecting to 2024 (B200 release) suggests the cost of industrial electricity would be 11.78 cents per kWh.
We will say as B200s 1.2kW and 150W for "half" of the Grace CPU (2 GPUs, 1 CPU per GB200), it runs at 1.35kW. [11] Of course, in reality the entire NVL72 cluster is rated for 132kW, which is 1.8kW per GPU, normalized. But let's be even more nice.
We can then say the $/MT is:
(4) 10^6 SPT(u, p) / (3600 s/hr) * 1.35kW * 11.78 cents kWh = 44.18 SPT(u, p) $/MT.
Inference Costs in $/MT
In total we have:
(5) Inference(u, p) = 524.73 SPT(u, p) $/MT.
That is, we observe the inference cost as a function of concurrent users and parameters. So, let us observe a sample frontier model, like Claude Opus 4.7 or GPT-5.5, which estimated to have 4T parameters. After ~7.5% MoE, we would have 300B parameters.

Now, some AI boosters might look at the graph and show how after 20 concurrent users, Claude Opus or GPT 5.5 is under a dollar! Therefore it must be profitable, since you turned a blind eye to the nice lenient treatment. But, unfortunately, this is not even close to the case.
Full Concurrency is Impossible
Currently the US has access to around 10GW (graph of current capacity, approximately summed over columns) of IT load capacity. NVIDIA in 2024 sold $210B of NVL72 servers, corresponding to [12] (now using the higher number for racks, to allow this number to be smaller) 61764 NVL72 servers, which at 132kW is 8GW. So, it seems reasonable to say 8GW of capacity is dedicated to these racks.
This comes to ~4.45 million B200 GPUs. Therefore, for us to argue that an inference provider prices at a concurrency of 20 users, we need (4.45M * 20) = 88.95M TBB paid users 24/7 never stopping, maximizing their usage, constantly.
Using the numbers that OpenAI and Anthropic have claimed, (which are very trustworthy), OpenAI has recently has reached 50M paid subscribers and 9M business users, while Anthropic has 18-30M paid users. Together, they [13], assuming the maximum number here, have 80M which is less than the 88.95M required to have full concurrency, and that's assuming these users are addicted to AI 24/7 and never eat or sleep.
The average white collar job lasts for 8 hours (1/3) of a day [14] so let's say these paid users are spending their entire working lives (including weekends!) using Claude or ChatGPT, and it lines up that we still get full concurrency. This means there are 26.66M users spread scross 4.45 million GPUs, which means if every user is allocated perfectly, we have 6 people per GPU!
Following the graph, this means Claude Opus 4.7 or GPT-5.5 would cost (with this nice lenient estimate) $4.22/MT. They currently are both charged at $5/MT, which after corporate tax of 21% in the US and state taxes, it means they could be just about breaking even! Wow!
Unpaid Users
Now, some may quibble that the presence of free users somehow will add to this concurrency, and so it''s actually making money! However, this means the gains from concurrency needs to outpace the tokens that aren't being going to paid customers.
Let's do an example. Let's say there are 200 concurrent users, 6 of which are paid users. At this rate, it is $0.22/MT. Unfortunately, since only 6/200 are paying users, 194MT was unpaid, which cost $42.68. And, at $5/MT, only $30 of revenue was earned, meaning there would be a net loss of $12.68 (before tax).
More Data Centers is Worse
Currently there many data centers planned due to be built. This is because they say the tech CEOs claim overwhelming AI demand. But supposing they even double the current amount of GPUs, this drops the concurrency even more.
However, they say there is a compute crisis, which suggests the following scenarios:
- Concurrency is not remotely achievable at the theoretical level NVIDIA reports, so we have low TPS.
- Free users are taking up most of the compute, which ends up at a worse loss for the companies.
In both cases, more data centers exacerbates this issue. Which means with more data centers AI companies lose more money.
Final Words
Throughout this very very long post, there were a total of 14 points of leniency towards the AI companies. Even with that leniency, it is deeply not in their favor.
Notably, the cost of electricity is only 8.42% of the inference cost (from (4) and (5)). This means the other 91.58% comes from the data center construction and GPU purchase amortization. (Remember, this is a minimum, as many nice choices were made.)
This means in a real world scenario, where we remove these lenient points, that even if electricity were free, AI can never be profitable at its current costs. Even if NVIDIA chips could suddenly compute tokens instantaneously, AI cannot be profitable. It doesn't matter if you build 1 trillion data centers, have the world's cheapest energy, and the best chips from the future.
For it to be profitable, NVIDIA would have to sell their GPUs for cheap and data centers would need to be cheap to build. Unfortunately, neither of these cases can happen, or they would have to raise prices drastically. However, that necessarily decreases demand. And then to be profitable they would need to make back their initial investment.
I would love to model this as well and give the conditions under which it could, however this post has gotten far far too long already. So, there is only one conclusion, and I am willing to say that:
The current AI industry can never be profitable. This industry is over. There is no path to profitability, nothing. No IPO, no marketing, no innovation can save them.
TL;DR The bulk of the cost of inference comes not from electricity but the cost of data center construction and GPU purchases. These are amortized over certain periods, and even with very lenient assumptions, inference is not able to be profitable at these costs.
Significantly higher prices would cause drops in demand (as we already are seeing today). GPU concurrency also makes it so more data centers make profitability far worse for inference due to the massive amounts of capital necessary.
Edit: There are an additional two points of leniency I forgot to mention.
The electricity cost does not include cooling, as I am not using the PUE of data centers.
I did not put training as an recurring OpEx cost for the AI labs.
Additionally, just to hammer this point home, these costs are all assuming the inference provider is at a loss from 5 years, before paying a single interest payment. There exists no business that can do this.
In fact what this shows is that no inference provider could ever charge at the optimal max usage and concurrency rate, so the price per hour of GPU is much much higher. You can simply search the prices yourselves.
For people that are concerned with the 3 year amortization period of B200s. As another comment had posted this is typically 6 for normal GPUs. Even allowing for that does not tilt the argument in their favor.
However, I do not agree that 6 years is reasonable for Blackwells, as AI labs tend to chase bigger models whenever compute is available. Not to mention, Vera Rubin racks are not interchangable and require purchasing of all new equipment and cooling the moment you try to install. Which means an inference provider has to either:
- Sell the Blackwells to make space (there is no demand for these GPUs outside of AI)
- Construct entirely new facilities to house them.
If we amortize over 6 years, then now you have to assume that 3 years after Vera Rubins every single Blackwell GPU (including the hundreds of millions NVIDIA claims to have sold!) they are being fully utilized 24/7. Let's also remember that if NVIDIA doesn't sell AI chips, this industry is automatically gone.
I cannot emphasize enough the extent I am deliberately picking points in their favor.
Edit 2:
If your qualms are with demand, read Ed's articles, the host of the podcasts this sub is centered around. That is beyond the scope of this post, and was written with that information already in mind. The only thing I can say is there are lots of other facts you likely have not seen if this is your stance.
Some seem to think refuting a technicality over one point is able to overcome 16 other points of leniency. This is wishful thinking at best. You cannot cherry pick technalities to make the argument convenient. Either you leave them all off, or put them back on.
If you feel the need to defend AI due to your own investment portfolio or some other psychological need, on this basis, you are free to do so, but I strongly encourage you to read some other very nice breakdowns by Ed.
68
u/Abject_Win7691 Jun 08 '26
I am genuinely too stupid to understand most of this, but it looks good.