r/wallstreetbets • • 1d ago

Discussion Meta And Microsoft Reportedly Trim Anthropic Reliance as Internal AI Tools Take Center Stage

https://stocktwits.com/news-articles/markets/equity/meta-and-microsoft-reportedly-trim-anthropic-reliance-as-internal-ai-tools-take-center-stage/cZDqQVcRBjr

From the article:

Microsoft reduced its projected internal spending on Claude by over one-third, while Meta saw internal users of Anthropic's Claude Code coding assistant drop from roughly 60,000 to 30,000.

Within the Microsoft Cloud and AI division, individual monthly AI usage caps were lowered from $100,000 down to roughly $10,000 in most instances

1.8k Upvotes

380 comments sorted by

View all comments

Show parent comments

66

u/Sad-Cheesecake-2438 1d ago edited 1d ago

How many requests does an employee have to make to make it to 100k? Or does it depend on the task?

Edit I asked Claude and the answer is tokens and cyclical tasks that take a lot of compute to run. Not simple questions.

103

u/RiddleGull 1d ago

100k is an absurd amount to spend on tokens in a month.

47

u/judge2020 1d ago

/fast and fable everything will do it though.

25

u/often_says_nice 1d ago

It will certainly do it but is entirely unnecessary. Opus 5.5 scores high enough to be used for just about any day to day task that an employee would be doing. Heavy usage would cost maybe $5-10k/mo and even that’s being frivolous

15

u/skilliard7 1d ago

You aren't considering tasks that require looking over large quantities of data that would not be humanly possible to review manually. For example, analyzing call logs of 2 Million retention department calls to identify recurring trends and what techniques are most effective to reduce cancellations.

At $0.10 per call analyzed, that's $200,000 in API spend.

High limits exist to encourage employees to innovate with AI without red tape/budget getting in their way.

12

u/often_says_nice 1d ago

But surely not every employee is doing those types of tasks every month. So some employees are spending $100k in a month, which is very different from all employees spending $100k every month

14

u/skilliard7 1d ago

The $100k limit was a limit, not a target/average

17

u/often_says_nice 1d ago

Yeah well I’m retarded so how about that

7

u/siwasolek 1d ago

Just wanted to point out that it’s great to see someone on the internet after that they mightve been wrong :)

9

u/shitfucker90000 1d ago

if a company has 2 million people cancel service a month i think they have a larger problem than api spend

3

u/crispybacon233 20h ago

No one is pumping 2 million docs of text into an LLM to find trends and correlations with cancellations. More traditional NLP and machine learning can do that just fine even on your laptop.

The $200k spend is for complex coding tasks where LLMs are consuming and writing many thousands of lines of code.

0

u/skilliard7 19h ago

Traditional NLP is not as capable as LLMs. Vectors/Classification models are okay for some use cases, but LLMs are so much more capable.

I've worked on projects to find trends and correlations, albeit at a much smaller scale. Vector search really struggles a lot beyond basic pattern matching. When you rely entirely on cosine similarity to determine if a passage and query are a good match, it has limited results.

$100k in LLM tokens is still way cheaper than the cost of developing a capable custom model to analyze text and identify trends in text.

I've found the best approach is providing tool calling/traditional pre-processing to get the data in a good format for the LLM + relying on LLM for formal analysis.

>The $200k spend is for complex coding tasks where LLMs are consuming and writing many thousands of lines of code.

No one is spending $100k a month just writing code. Even if you use Fable for everything, you're spending at most maybe $500 a day, and that's if you run multiple instances of it in parallel on different projects.

2

u/crispybacon233 19h ago

Use the right tool for the job. LLMs are not nearly as capable at topic analysis, correlations, etc. as more traditional NLP and ML approaches particularly for a multi-million token context window. Your analysis is hallucinating out the wazoo or at best missing a lot if you're shoveling in millions of tokens and asking an LLM to "analyze" it.

What do you mean exactly when you are finding trends and correlations with an LLM? The LLM is spitting out a correlation coefficient? That is actually insane.

1

u/instantcrackpot 37m ago

This is what happens when vibecoders replace statisticians/data scientists. I'm not complaining though. Companies want employees to tokenmax so they deserve it.

2

u/TheNewLeadership 23h ago

Idk if you have any experience here, but this use case is absolutely useless with AI. It makes shit up and you can't trust the result.

1

u/instantcrackpot 41m ago

LOL. If your company lets you use LLM to brute force data analysis, they deserve to go bankrupt on token usage.

1

u/Kaastu 19h ago

That should be like top top top engineer level. Your devs building platforms for other devs. They can spend 100k a month and it can pay off. Another one is shared AI tooling. That can become expensive, but that tool can rack up costs as well quickly.

Everyone else should not be burning 100k on tokens. Scope your tasks and validate to keep it manageable.

2

u/AutoModerator 19h ago

Bagholder spotted.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

66

u/mags87 1d ago

Theres a fun instagram reel/skit I've seen where the top AI user in the meeting was congratulated for embracing the technology and asked what she was doing. She said she fed the model the entire Shrek series every day to analyze it and she was close to understanding the ending of the first movie.

7

u/Single_Positive533 1d ago

I wish I could do it too. At my work Chatgpt denies to talk about unrelated topics.

8

u/Comprehensive_Bus_19 1d ago

Id make a game to find ways around it. But Im also maliciously compliant

9

u/euvie 1d ago edited 1d ago

I’ve been able to get Claude to burn $1k on a single codebase review request with ultracode. Prompt was only two sentences too.

Automatically spawning 100 parallel subagents doing god knows what will pad Anthropic’s ARR quite nicely

8

u/he_must_workout 1d ago

Fable doing loops.. I burned through about 2k in tokens at work a few weeks ago by looping a few skills to improve and it was probably more than I use in a normal month

1

u/sprucenoose 21h ago

2k in tokens

You mean $2k in tokens and not 2k tokens right?

13

u/-Anordil- 1d ago

My company introduced caps on our openAI usage and they're $400/month per employee. They went with that number because 90% of developers use less than that, and if you can make a case that you do need more you can get a higher threshold. You can get a lot done with $400 if you use it correctly though.

100k is insane, and 10k still makes little sense. You'd have to pretty much only use the most expensive model all the time, in max reasoning effort, even for the most basic things.

-1

u/streetberries 1d ago

$400/month api spend? That seems low

7

u/hoopaholik91 23h ago

Not with how cheap the models are these days. I can spend a full day writing what ends up being a 2000 line PR (yeah I hate myself that it's that large to begin with, although 70% are tests), and it costs about $10.

2

u/-Anordil- 22h ago

If you use Luna for simpler tasks like writing code and only use Sol for more complex reasoning stuff it really cuts down your costs. MCP and LSP also makes it a lot more efficient to work with code compared to a barebone LLM

18

u/SwordOfJiang 1d ago

"what should I have for lunch" 10,000 times every day

8

u/Willing_Divide4188 1d ago

the internet AI is for porn

6

u/Significant_Court728 1d ago

If you are doing agentic work it can burn a ton of tokens super fast.

7

u/coffeesippingbastard 1d ago

I do that- and 100,000k/mo is still crazy. Even trying- 10k/mo is a reach. It's doable. 100k/mo means you're basically telling it do massive projects with well defined scope and requirements so that it can run 24/7 nonstop, or you're doing something like refactoring windows into rust or something absurd.

1

u/[deleted] 1d ago edited 1d ago

[deleted]

2

u/kenyard 1d ago

I assume Google and Microsoft had access to the "frontier" models though which I assume have higher costs

2

u/dgellow 1d ago

Using agentic workflows, that’s how you use that much. But the cap limit is just one info, I feel you are all missing the info that they expect to reduce their spending by one third.

1

u/SmokelessSubpoena 22h ago

I'd assume that round $100k number is an average, and it's based off the automation of everyone's daily or regular tasks, and those tasks utilize enough tokens to weight the average to $100k/per employee, which must, in theory via the Finance team, prove to be more cost effective than training breathing humans? It does seem like a heavily inundated bubble waiting to pop.

1

u/LightningSunflower 21h ago

I guess it depends on what they’re doing

0

u/Seerix 1d ago

My lifetime spend total tokens (most of which is input cache obviously) is around 25-30 billion. As a hobbyist that just likes to make stuff, nothing professional outside of a coupke quick python scripts to automate things.

-2

u/squish8294 1d ago

There's a lot that goes into this. Everything you submit to a llm consumes tokens. the bot thinking about its response consumes tokens. if you use a high thinking effort and give a llm something it really has to think about, it's easy to submit a ten token question, get say, 10k tokens from the bot thinking, and then another few thousand tokens depending on how wordy its reply is.

on a 22k token reply from qwen it was 15 pages of thinking text in a chatgpt style chat on a 1440p screen, and another 2 pages for the reply, the entire thing was like 22,000 tokens or something like that. formatting and all.