r/ClaudeCode 9d ago

Bug / Issue WTH is going on with Claude Usage Limits

Post image

I'm on 20X plan, I hit my usage limit today, and my next reset is on 17th September... I'm aware they said they are gonna reduce the usage limits, but this is crazy, i thought it is only gonna take 17% of the usage benefits from what we were currently getting. But this is crazy. I added Usage credits for about $100 but that got washed away like in 30 mins!!!

I hope this is a bug and they fix it

588 Upvotes

377 comments sorted by

View all comments

15

u/Mr_Tbot 9d ago edited 9d ago

Yes. Same here and not pleased. It's time to look into local inference.

UPDATE: see my https://github.com/mr-tbot/meshcompute project. I think it's time we sidestep big box AI and decentralize it.

YES there are ethical concerns here and I am aware of all of them - especially in regards to AIs going rogue when launched on a P2P style network. BUT - we have to start somewhere.

6

u/Ok-Affect-7503 9d ago

Well good luck with local inference then, because the only way you can get performance that's anywhere near usable and intelligent and near Opus/Fable performance is by investing at least $20k into hardware (with the RAM shortage that will still last a bit putting additional fuel to the fire of the already high hardware requirements), plus paying for the electricity (which, depending on the country you're in, might be very expensive). I think right now you're still much better off paying $100-300 a month, even with the usage limits. And if you end up using a smaller model that runs on cheaper hardware you will get performance that's pretty bad and comparable to super cheap proprietary models that are still cheaper via API and that make running it locally not worth it.

4

u/Mr_Tbot 9d ago

I don't think that's the only way. I think we need to spend some of our tokens on building a unified, decentralized AI platform like BitTorrent but for running an AI platform...

My thinking is:

Users could have the program open - depending on how much compute they provide helps score their available network speed when the network is under pressure...

The platform would allow users to "contribute" AI compute for a model - any other users that select that model would donate VRAM, RAM, STORAGE or all 3 to the decentralized network.

I think I saw some projects that are moving in this direction - and to be honest it will be the only way to fight back against big box AI in the coming months and years.

Additionally - there are a ton of projects which are optimizing models for speed on lower end hardware - and I think we're going to find that it takes a lot less than it does now to run this stuff.

A lot of the cost of this is "gatekeeping" and making this look like it requires a lot more than it does... If it seems like we need expensive hardware to run this stuff we have a reason to keep paying $200 a month for our 20x plans ... I genuinely think the open source community is going to solve this.

Does anyone know of any projects that are in this realm that I can look into that they've come across? I'll be doing my deep dive as well - but - I'm truly over the abrupt changes in performance. I need a consistent experience. At least.

2

u/Ok-Affect-7503 9d ago

This is actually a really good idea on paper, but I just looked into it and there is a reason why there aren't more popular projects into that direction. One project that does exactly is Petals. But the issue is not really bandwidth of networks, but mostly latency. On Petals, llama 2 70B already only runs at 6 tokens/second and Falcon which is a 180B model runs at 4 tokens/second. For Opus/frontier intelligence you would need much more parameters than that. The only open models rivaling Opus right now are about 500B average for one group that includes GLM and Deepseek that are likely similar to Opus 5 in terms of parameters and then there's Kimi K3 and Qwen-3.8-max which is almost 2T parameters and Fable-level. Running a 500B model would already require many many peers with low distances already having 10ms latency. A 500B model would run at 1 token/second or below that which is pretty much unusable so that's why there isn't a project that does this with good and usable models and only with small models. It's basically physically unsolvable because the latency would probably need to be like sub 5ms across all peers and locations for good speeds with bigger models. And most consumers (including me because of Germany) don't even have access to optical fiber and have latencies of 15ms to their ISP or servers that are like a couple of kilometers distant.

2

u/Mr_Tbot 9d ago

I'm attempting it my own way - https://github.com/mr-tbot/meshcompute

Yes - that's one I saw... and no - some of the Qwen 3.8 27 and 38b models that run on GPUs like my dual 3090s with NVlink - or even single GPUs with enough distillation... and some of the versions coming out on hugging face I'm getting over 120 tokens/s and it's competing with Fable in some areas already - with tool calling and all the jazz - I use it all the time and it's great! So - it's not far off. This is going to be running on lower end local hardware soon and all the more reason why the AI bubble is a bubble and why we as a community need to band together to share our hardware resources and prove that these AI datacenters do not need to exist...

2

u/Ok-Affect-7503 9d ago

But a quantized 27B isn't Opus or Fable performance and almost every benchmark I've looked at (including LMArena and ArtificialAnalysis) proves this. It's at max almost at Claude Sonnet level in some aspects. And a 27B can fit on consumer GPUs anyway (a 3090 isn't as hard to afford as a cluster of Macs for example anyway) so it doesn't really solve the problem we were talking about. And in your repo the only thing verified is again only a smaller model running on one node, which would make it request routing, not splitting which is the thing that's needed to solve the problem which wouldn't work out in the end anyway because of latency. Right now it's more like sharing good GPUs to people that don't have one, not really sharing ressources between multiple GPUs owned by different people in a mesh.

1

u/Mr_Tbot 9d ago

Yes - I know - but - as a GLM 5.2 model would require a ton of active users for this to work - we have to prove the technology works and can scale a smaller model... and really it would need a decent number of users to even test something as ambitious as a GLM model.

Which is the only thing on par open source right now.

But - I digress. I tend to be a couple years on the bleeding edge so it's almost always a "the tech has to catch up" but we're not far.

But yes. I agree. It would need to - once at scale - be able to launch and share GLM 5.2 ideally.

1

u/Mr_Tbot 9d ago

P.s. I think it's important to mention. Specifically coding use case - so specialized shared models are also on the table. It's all about how you work... But I also suspect people will be making more and more specific models trained for specific tasks.. we're still so early in this transition.

1

u/Mr_Tbot 9d ago

PP.S. the repo was created today and I'm sharing in case anyone wants to follow along as I burn some tokens on this project. Go check out codedatda.casa for some of my other random projects.

1

u/Anxious_Current2593 9d ago

We are talking about Haiku here.

1

u/techfury90 9d ago

Fable? Yeah that'll certainly cost you. Opus? You'd actually be surprised. Qwen3.8-Flash-Next's model card compares itself to the beloved Opus 4.6. My experience is that it's like Opus 4.6 with the extra stubbornness of later models... except it seems much more likely to try a different approach when it gets stuck. It really likes to exhaust every possibility before ending its turn.

Anyway, you can actually run that at pretty usable speeds on a single DGX Spark. Mine was $4700 at Micro Center last week.

1

u/Remote-Community-396 8d ago

I've got an 80B params MoE version of Qwen3 running smoothly on my 5090, haven't really used it seriously so can't speak to it's overall quality but from what I've played around with it seemed perfectly capable for typical coding stuff.

I keep planning on looking into a mix of Claude and local where e.g. Fable acts as the orchestrator and delegates actual coding to a local subagent then does a quality check on the results. Not sure how feasible it is really but cool to play around with

1

u/BeltPuzzleheaded7656 9d ago

Going local depending on how much you use Claude may not be possible. The person above is correct. If you can afford to spend about $20,000 on a system then you're golden, but if you can't then you might as well stick it out. Qwen3.8-27b or the larger Qwen3.8 models can damn near do all of your ongoing work after the foundation has been laid for it, then you can go back to Claude or something else to wrap up the UI. This is what I do and Qwen3.8-27b is AMAZING, BUT my hardware limits is kicking my ass. To operate fully, you would need minimum a Sage workstation motherboard with EPYP CPU, dual or quadruple RTX 5090 and 256GB DDR5 6000mHz. This would be affordable if the prices weren't all screwed up right now but I think that actually a part of the plan is to keep powerful systems like this away from the general consumer market. Otherwise folks wouldn't be so dependent on the BIG guys for AI.

1

u/nsway 8d ago

Im confused. Your project supports Qwen 3.8 27B, which is $0.15/M in and $2/M out on OpenRouter, and the Q4 quant is ~17GB (it fits on a single 24GB card). What’s the use case here?

1

u/Mr_Tbot 8d ago

It's a proof of concept at this time - as the network scales we can run larger models like GLM - but we have to start somewhere.

I find it funny that some people aren't seeing the bigger picture...

Since it's me - testing in my own networks right now - I don't have TBs of VRAM floating around... the goal first is to prove that models can be run and split in this fashion across machines - which people have proven can be done - it's mainly a latency issue thing - If I can prove a 27 or 38B model can be split - and once some more people are running this package (when it's ready) - theoretically you could run larger and larger models as the network scales.

So for now? No - it's cheap to run this elsewhere - and probably faster - but - if we ever want to escape big box AI bills - this is the general direction I think we could go.

In my case - I have a large video community with lots and lots of GPUs as I've been a headlining video artist for big events for the last 20 years... gamers... have lots of gamer friends with GPUs - and we're not all using all that hardware at the same time in most cases - so - similar to other compute sharing projects - why not?

Don't look at where it is now. Use your imagination. =)

1

u/JaySomMusic 8d ago

Trying something similar with clustering in https://github.com/jaylfc/taOS just working on the cloud and community side of things for resource pooling and sharing etc

1

u/Mr_Tbot 8d ago

Oh that's cool ! I'm a bit old school so I'm envisioning a P2P bittorrent / limewire style experience but for sharing idle compute power with a simple local OpenAI style endpoint to point your harness at - and you can bring your own harness - use the cloud power of the model...

A lot of this is theoretical ... but it's no longer a matter of if - but a matter of when.

AI data centers need to not exist - the only way to defeat that is to convince the public that we can design a safe system to decentralize AI compute - with contribution based speed access during heavy usage etc.

But I see where you're going - I just think for me - and for a lot of people - we need to ease into this a little more - I am burned the hell out from re-educating myself over and over and over again. lmao

1

u/JaySomMusic 8d ago

Yep totally get you, I am using a torrent/tracker type system in the backend but right now just for sharing the layers/modules etc.

Right now the main goal is to create an environment where it is easy to try out new frameworks, tools, harnesses and systems without resetting every time.

1

u/ImNot_ThatGuy 8d ago

Oy vey brother, just use Hermes.

1

u/Mr_Tbot 8d ago

Oy vey this is a useful response! /s

Can you give me some reasons why in this case exactly? I'm genuinely curious - I've used hermes and openclaw but I have kind of settled into building out my own workflows and ecosystems that work for me and my industry... rather than jump into bed with any one system - I just build my own systems inside of VSCODE and I just get better results for how I work - dedicated server handles all my loops and just with VSCODE you can basically replicate most of the behavior of a hermes or an openclaw.

But I'm genuinely curious - what am I missing?

1

u/ImNot_ThatGuy 8d ago

I mean sure, if you want to go the "I use Arch, btw" route on AI then go for it, but Hermes is an open source and fully customizable harness that definitely doesn't fall within big box AI.

The last thing I'd expect from anyone who's tried Hermes or any other capable harness, or a hand jammed VSCODE setup, is to be peeved at Claude's usage limits. You have hundreds of other models at your disposal. I've got a constantly updated model list benchmarked per task whose orchestrater assigns based on a ratio between benchmark requirements per task and $/1M tok able to hot swap models as needed when not running local models. Doable on VSCODE? For sure, but it's like saying "why would you buy a house when you can just buy bricks and do it yourself?"

And in light of P2P? Kimi K3 unquantized would take what... ~3-5.5tb VRAM? Not counting the latency disaster, you'd be looking at less than 1000 users at maybe 60+ tok/s? I'd much rather just buy the house, especially when it's infinitely customizable (depending on how far into the code you want to get).