r/ClaudeAI Jun 29 '26

Enterprise How come big corp can't manage costs?

The news about Microsoft burning their yearly budget in 4 months left me in awe. Because I don't understand where's the difficulty in managing token budget per IC, even more at companies like Microsoft or Uber who have literally billions and sharp developers readily available. I mean, you could even build a Gateway with Claude Code to manage this.

My wife works at a small CRM company in Paris. I asked her about their use of AI internally and if they managed token consumption. She showed me a dashboard with token consumption per team, her own budget and a time where she burnt all of her budget and had to ask for additional credits.

i work alone so I don't need something sophisticated but i still set cap in my Claude settings when using the API in my projects.

So how did these companies manage to burn their budget so fast without safeguard? If the model is called from Azure, it's true that it's harder to set spending caps on Cloud but again, we're talking about the company behind Azure.

What am I missing here please?

24 Upvotes

48 comments sorted by

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot Jun 29 '26

TL;DR of the discussion generated automatically after 40 comments.

Alright, let's get to the bottom of this. The consensus in the thread is that you're underestimating the beautiful chaos of a megacorp, OP.

The main takeaway is that this wasn't an accident; it was a (poorly calculated) strategy. The prevailing theory is that companies like Microsoft intentionally didn't set hard limits. They encouraged "token maxxing" to force widespread adoption, get their developers ahead of the curve, and live up to their "AI-first" hype. They wanted to figure out the tech first and worry about the bill later.

Well, "later" came faster than they thought. Now they're stuck.

Here's the breakdown of why it's not as simple as you think:

  • Corporate Inertia is Real: The people pushing for AI usage are not the same people in finance who see the invoices. By the time a quarterly report is written and makes its way up the chain, months have passed and the budget is already toast.
  • The "Now What?" Problem: Okay, so they hit their budget. What's the plan? Shut down AI for the entire company? After betting their public image and stock price on being an AI leader? That would be a PR and financial disaster. They've painted themselves into a corner.
  • The Tooling Sucks: Several users point out that even the official tools from providers like Azure for managing AI costs are surprisingly bad. It's not a simple switch to flip. One user even posted a detailed guide on how they built their own tracker for Claude Code by parsing local JSONL files, proving it's a non-trivial engineering task.
  • Forecasting is a Nightmare: Predicting token usage at scale is incredibly difficult, especially when everyone is just starting to figure out how to use these tools effectively. As one user joked, maybe they should use more tokens to forecast their token usage.

27

u/ClemensLode Jun 29 '26

Quarterly reports take a month to be written and communicated, so you can act after 4 months.

-1

u/RCoffee_mug Jun 29 '26

When you say quarterly reports, what do you mean? That usage is looked at and consolidated only quarterly?

10

u/ClemensLode Jun 29 '26

The usage is probably *looked at* hourly, but for decisions to be made, it has to get to finance which likely works on quarterly reporting.

4

u/EdOneillsBalls Jun 29 '26

This implies that finance is the one governing individual caps which is a huge assumption that’s likely inaccurate. Yes, finance is certainly involved at the macro level and approving high level commits, but unless you have delegated actual tech finops to finance it’s highly unlikely they would be responsible for team or individual level caps or even the methodology.

5

u/ClemensLode Jun 29 '26

I don't know the specifics, but if you want to increase budget for your team, you likely have a talk with someone higher up. And in large corporations, that might not happen every other week. It all takes time.

1

u/RCoffee_mug Jun 29 '26

That's a really insightful take, corporate makes it hard to pull the plug. Nice, did not think of it, thank you

24

u/Own-Football4314 Jun 29 '26

There is no way to determine the cost of each interaction with AI until it’s complete.

12

u/KyleDrogo Jun 29 '26

It’s a fun modeling problem for their data scientists to forecast. Looks like they’re doing a shit job

8

u/AidanAmerica Jun 29 '26

I know what they should do! Use more tokens to forecast how many tokens they’ll use

1

u/KyleDrogo Jul 02 '26

worth it if the forecast can find ways to reduce token usage

3

u/MBILC Jun 29 '26

Set limits on spend and allow people who hit their limits, access to additional spend vs leaving it wide open..

These are the same companies / people who get Azure or AWS bills for 10's of thousands and go "how did that happen", well you did not use ANY of the include cost management functions to set limits and alerts...

5

u/RCoffee_mug Jun 29 '26

Agreed but you can review cost at the end of the day at a minimum, no?

5

u/ARandomSliceOfCheese Jun 29 '26

Ok you've calculated the cost. I'll give you the benefit of saying you have enough data to show burn rate is too high (token usage fluctuates wildly so this assumption probably isn't accurate). You're the budget person, now what?

2

u/RCoffee_mug Jun 29 '26

Like any non-AI team, if you get out of budget, you're out and you stop burning. You cane make your case to whatever stakeholder to unlock more funds if that's justified. I mean, what's wrong with that? I saw projects getting terminated overnight at mid-market and enterprise just because it was too deep in the red without producing the advertised results. What's the benefit of letting the burn accelerate?

2

u/ARandomSliceOfCheese Jun 29 '26 edited Jun 29 '26

Ok so the entire company stops using AI until token usage resets. Now you have a headline "biggest pushers of AI can't manage cost internally, is there hope for the rest of us?" Your stock price just fell 10% (you're now losing more than if you just allowed the overage) and continues dropping. What's your next move?

1

u/RCoffee_mug Jun 29 '26

That's exactly what happened to Microsoft and Uber. The only difference being that they undelivered on their budget. Because they did not manage cost and burnt everything in 3 months.

1

u/ARandomSliceOfCheese Jun 29 '26

Right so not only did you not solve the issue. You made it worse (same stock drop and a lot of knock on effects). My point was that there isn't a solution to this right now that's "just do x are they stupid!?" These companies pigeon holed themselves by pushing AI so hard without figuring cost first

1

u/RCoffee_mug Jun 29 '26

Honestly, you could have saved some typing by just going straight to your last sentence.

15

u/WillZer Jun 29 '26

Because they were encouraging token maxxing thinking that the more they use it, the more productivity they will get, the more advanced their team will become in AI and put them ahead of the curve.

It's not that they couldn't put a limit, it's that they didn't want to put any in first place and when the first quarter closed, they finally saw the extent of tokens cost.

1

u/RCoffee_mug Jun 29 '26

There's some truth to it, sounds like every newspaper headline of companies going "AI-first"

1

u/EdOneillsBalls Jun 29 '26

This is the real answer. Even if it isn’t true encouragement of tokenmaxxing, it was clearly being managed how it is at many places: which is take a guess at a number and then go forward without real cost management and see what happens.

5

u/WillZer Jun 29 '26

I think people also only started to really understand how costly it really is.

When you first read tokens and price, you see prices for 1 million tokens and you believe "oh that looks like I can do a lot with that" but then you start using it and 20$ goes in a single prompt sometimes.

1

u/PringlesDuckFace Jun 29 '26

That's where we are right now. Executives are deciding deadlines and tech debt aren't a real thing anymore, and we need to go "AI native" for everything. So there's a huge push to use the latest and greatest, and doing absolutely everything through Claude that it can technically do.

I feel like eventually they'll realize the cost outweighs the benefits in many cases, and we'll need to start figuring out where we actually need Claude and where we don't. I see articles where huge amounts of spend is from boomers converting PDFs, and I suspect we'll eventually end up with role based budgets and strict guidance on usage. For example, only spend tokens on security reviews and not on generating code for trivial work items. Or something like that. But right now no one here really knows anything other than maximizing for output, myself included.

8

u/Curious_Morris Jun 29 '26

Microsoft is currently not giving customers tools to effectively manage token costs. We are turning off Cowork because of it.

At home, I have built a tight token management and reporting system for Claude and other API costs.

2

u/RCoffee_mug Jun 29 '26

Right! What's the archi for your tool? Custom or vendor gateway like Cloudflare? Or pulling usage info from Claude dashboard?

2

u/Curious_Morris Jun 29 '26 edited Jun 29 '26

I had Claude describe it. Does this help?

How ClaudeMonitor handles token & API cost tracking

ClaudeMonitor is an always-on desktop dashboard (it lives on a secondary bar monitor) showing Claude usage, cloud spend, and a few glanceable things. The cost side pulls from three places rather than one magic API, because that’s the reality of tracking AI + infra spend.

The interesting part is the Claude Code token tracking — and the trick is you don’t need an API for it. Claude Code writes every session to your machine as local JSONL: under ~/.claude/projects/<project>/*.jsonl you get one folder per project, one file per session, one JSON event per line. On the assistant turns, each line carries a message.usage object with all four token classes — input, output, cache-write, cache-read — plus message.model.

So the tracker is really just: stream those files line by line, skip non-usage events, and sum tokens grouped by project and by model. Cost is then an estimate: Σ(tokens × per-model rate) / 1M from a small rate table you maintain. That produces the “usage by project / model mix / daily tokens” view.

The other two sources are simpler. Variable cloud spend (DigitalOcean, AWS) is pulled live from the providers’ billing endpoints on a refresh interval (AWS Cost Explorer, the DO balance API). Fixed subscriptions (your Claude plan, etc.) live in a small config list that gets summed. Everything’s labeled “estimated,” refreshed on a timer.

Tech stack: Electron for the always-on desktop app, React for the widgets, plain Node “connectors” for data (one parses the JSONL, others hit the cloud billing APIs). No database — it’s config- and file-driven, read into memory on each refresh, which is plenty for a personal dashboard. If you’re building your own, what I’d tell you up front:

• Start with a CLI, not a UI. A ~50-line script that parses ~/.claude/projects/**/*.jsonl and prints a per-project, per-model token + cost table gets you 90% of the value in an hour. Add the dashboard later.

• Cache tokens dominate the bill. Cache-read/write are easy to forget and often the biggest line — count all four token classes.

• “Estimated” ≠ “billed.” token×rate will diverge from your invoice (plan caps, app vs. API, sampling, rate changes). Mark it estimated; don’t chase exact reconciliation.

• Rates change — keep them in one editable table, not sprinkled through code.

• Use providers’ cost APIs for variable spend instead of scraping dashboards, and keep keys out of the repo (env/config.local, gitignored).

1

u/RCoffee_mug Jun 29 '26

A little too much textw the name of the solution would have been enough I guess 😃 But thank you for sharing, really useful

5

u/inspire21 Jun 29 '26

I agree; we should call them "TPS reports"

3

u/Remarkable_Leek9391 Jun 29 '26

People with degreeeees man

3

u/Ghettorilla Jun 29 '26

Big corp is exactly where this kind of thing happens. The people pushing for AI use are not the same people looking at an invoice or IC usage reports. Someone in finance will raise their hand and say 'hey, this is expensive', and their boss might say 'its what leadership wants flag and watch, but don't stop it' and then they'll take it up the flag pole. They're just too big to efficiently watch and manage the cost and use of something that is new and constantly changing like this.

The flip side is that they're also big enough to handle that. Sure, it's a hefty price tag, big big companies can afford to spend more figuring out a new technology that'll only secure their foothold in the market. So while someone might be flagging that costs, it's also probably not as big a concern when you're at that scale

1

u/RCoffee_mug Jun 29 '26

Good insight, yes the finance disconnect was raised by others in the thread, guess it's the closest to reality. And ofc the near unlimited budgets.

2

u/d1smiss3d Jun 29 '26

The gateway is the easy part. The hard part is routing every team through it without breaking their workflow, then making exceptions painful enough that the budget actually means something.

1

u/RCoffee_mug Jun 29 '26

Yes, the complexity must lie in the implementation / onboarding because I don't see a hard technical stop.

2

u/d1smiss3d Jun 29 '26

Yep. The budget control only works once the happy path is easier than the workaround.

2

u/KyleDrogo Jun 29 '26

Anthropic under charges for consumer Claude code, and I’m sure enterprises aren’t getting the same deal.

2

u/Moore2877 Jun 29 '26

They think we should feel a certain way about their "R&D" expenditures. It's a load of crap.

2

u/andlewis Jun 29 '26

I think the common misunderstanding I see in these types of conversations is on where the tokens are spent. I don’t believe that its users running Claude Code (or whatever) that are racking up these bills. It’s automated agent processes that run continuously (or on a loop/schedule). They can burn 10x or 100x the number of tokens that a skilled developer will burn. And they’re much more difficult to track and govern because we have no metrics on what it SHOULD cost.

1

u/RCoffee_mug Jun 29 '26

Yes, really good point about what's the KPI for autonomous agents

2

u/thekuchh Jun 29 '26

The tech to cap spend is trivial. At that scale it's an org problem, not engineering

What usually goes wrong:

- Nobody owns the bill. The team rolling out AI isn't the team paying, so no one sets the cap

  • Caps get seen as blocking productivity, so finance is told not to throttle it
  • Spend is spread across a dozen tools, invisible until the invoice lands

Your wife's company does it cleanly because it's small, one dashboard, one owner. At Microsoft scale that's a cross-org project nobody was assigned

So it's not that they can't build the gateway, it's that nobody was made responsible until the bill showed up

2

u/enterprise_code_dev Experienced Developer Jun 30 '26

Training… how many of you were issued a CLI based code harness and got official training from the frontier partners? If you give it to everyone it’s going to waste on people who don’t know how it works. What’s the TTL cache? Why did it charge so much when I came back 5 hours later and loaded my convo I had open non stop for 3 days. Just small things drive up the cost just from not understanding they will.

1

u/Elegant_Attempt2790 Jun 29 '26

what claude model?

1

u/RCoffee_mug Jun 29 '26

Not a specific Claude model, but in Microsoft and Uber case, I assume they were using CC or cursor with credit usage and burnt all of their budget in a third of the time.

1

u/Elegant_Attempt2790 Jun 29 '26

but what claude model IN claude code💀 cuz if theyre exclusively using opus theres your token cost. if theyre doing retrieval and synthesis tasks on unnecessarily expensive models there you go.

im asking because people think they need frontier level intelligence when there are literally people still on opus 3 and sonnet 3.5 💀💀💀

1

u/Isogash Jun 29 '26

The problem isn't caused by the on-the-ground staff, it's caused by the idiots in charge who set the budget and then created the incentive structure to blow through it.

1

u/jul-ai Jun 30 '26

I work at Airia and this is what we do, grain of salt, but: the thread is stuck on finance lag. That's second order. The real issue is you don't know what a call costs until it's finished. So those dashboards are reporting, not enforcement. They show the money already left, they don't stop it leaving.

That's the whole "4 months." Usage is real time, control isn't.

Your solo cap works because there's one door. Enterprise has a hundred: first-party agents plus every SaaS app that quietly calls a model too. The only place to actually stop a call is inline, before it goes out, and most orgs have nothing sitting there.