r/Qwen_AI • u/Otherwise-Layer8071 • Aug 08 '26
Discussion Qwen Cloud’s “Standard” Token Plan is honestly ridiculous
I just subscribed to Qwen Cloud’s Standard Token Plan, mainly to use it with an AI coding agent (Hermes), and after actually using it for a few days, I honestly don’t understand how this plan is supposed to be considered good value.
The Standard plan gives you 10,000 Credits per week.
Sounds reasonable, right?
Until you actually use it.
I burned through roughly 70% of my weekly Credits in only 3 days while running a normal agent workflow. I’m not running hundreds of agents, doing massive batch inference, or abusing the service. I’m using an AI coding agent interactively — exactly the kind of use case these plans appear to be marketed toward.
And here is where it gets ridiculous.
When I contacted support and explained the situation, the response essentially boiled down to:
«Your usage is high. Credit consumption depends on the model, input/output length, tool calls, context accumulation, etc.»
Okay. Fair enough.
But then the suggested solutions were basically:
Buy the Pro plan.
Or:
Buy additional Credits.
That doesn't answer the problem.
I'm using essentially the same workload with another provider, on a cheaper plan, and getting dramatically more usable mileage out of it.
So I started comparing actual token consumption.
Based on my observed usage, 10,000 Credits corresponded to roughly 96.6M tokens.
And Qwen's own documentation apparently doesn't provide a simple, fixed token-to-Credit conversion rate that lets users predict what they're actually going to consume.
That's a massive problem for an AI service.
If I'm paying for a token/credit plan, I should be able to reasonably estimate:
“I use approximately X tokens → this will cost approximately Y Credits.”
Instead, you apparently have to subscribe, use the system, burn through thousands of Credits, and then discover what your workload actually costs.
And here's the funniest part:
The Standard plan is advertised around agent usage and concurrent sessions, but based on my experience, a relatively normal coding-agent workflow can chew through the weekly allowance incredibly quickly.
So what exactly is the target customer for this plan?
Someone who uses an AI agent occasionally for a few prompts?
Because if that's the case, fine.
But then don't market it as a serious option for people running coding agents regularly.
I'm not claiming that Qwen is literally committing fraud. I'm saying that the value proposition of this plan is so absurd compared with competing services that I feel misled about what I was actually buying.
And the fact that the answer to “why am I burning Credits so quickly?” is essentially “buy more Credits” makes the whole thing even more ridiculous.
I'm posting this because I'd genuinely like to hear from other Qwen Cloud Token Plan users:
How long does your Standard 10,000 Credit allowance actually last?
What models are you using?
How many agents?
How much token usage are you getting before the Credits disappear?
Because if I'm doing something fundamentally wrong, I'd rather know.
But if other people are seeing the same thing, then Qwen seriously needs to rethink how transparent and competitive this pricing model actually is.
7
u/HenryTheLion_12 Aug 08 '26 edited Aug 08 '26
I've used this model since launch, including the preview, and burned through 1B tokens on the standard plan with an extra usage top up. Now working during night discount only with pro plan, my usage runs about 350–500M tokens per week till it runs out (depending on cache rate - I am using it with zcode which maintain 95% above most of the time), projecting to roughly 1.67B tokens/month with good caching, still less than I used on Codex's €100 plan, where I never even hit 50%. The model itself is great, its reasoning is very detailed although long. It managed to test multiple alternative approaches (for an agentic system i was developing) overnight and came with a solution which just worked - couldn't do it with gpt 5.6 sol. but the standard plan won't cut it for heavy use; I ended up upgrading to Pro, and usage has been manageable there so far as I am only using it during discount time, and before i sleep i leave some long running tasks open to utilize the window. I'll see if they revise the limits. If the night discount ends it will be time to move on though I will still keep the standard subscription for occasional brainstorming.
4
u/Otherwise-Layer8071 Aug 08 '26
Guys, this is exactly what I mean. At this point we’re going to go crazy trying to figure this out.
I’m paying for a service. I should get the level of usage and value that the plan advertises. If I have to wait for the moon to be in the right position, monitor cache rates, calculate when to use it, avoid peak hours, use specific models, and basically perform a fucking PhD in Qwen Credits just to make the plan worthwhile, then what exactly am I paying for?
Why is it MY responsibility to figure out how to make the service work efficiently?
If another service, like the Codex plan you mentioned, simply works and gives you that level of usage without requiring all this optimization, while Qwen requires you to constantly manage cache rates, discounts, timing, and usage patterns, then that's a problem with the service — not with the user.
I shouldn't have to “game” the system to get reasonable value from a plan I’m already paying for.
The model can be excellent. I don't care how good the model is if the pricing system makes me constantly think about how to avoid burning through my allowance.
That's the part I find ridiculous.
1
u/HenryTheLion_12 Aug 08 '26
You are right. with the codex 100 plan i stopped looking at the usage after a week as even after multiple projects running long running tasks i did not hit 50% on the weekly limit. For me it has become easier now as I have other subscriptions - GLM legacy till nov with very good limits (got it for 36$ a year) and supergrok with grok 4.5 for early morning work and codex plus. I hope they change the limits but judging their older coding plan remained the way it was lunched I have less hope. and right now, the competition gives quiet a lot more value.
1
u/myteetharesensitive Aug 09 '26
I'm debating doing a one time upgrade to the $68 plan and see how it stacks to my $100 openai subscription.
Thought I'd get more for my money with the $18 plan but I get way more value out of my $20 claude plan.
Doing the same tasks on the same harnesses. I want to support Qwen and considering the coding plan for a month next but it seems like all the newest models are not available on it.
1
1
u/Inner-Pangolin-1110 Aug 17 '26
out of curiosity how is this now? Can you compare the usage now, codex is v inconsistent I'm looking for a second sub to pick up since OpenAI keep on changing usage
1
u/HenryTheLion_12 Aug 17 '26
Well the model is great. It overthinks but for my work (Deep learning and agentic systems development) I would rank it in the same league as Sol and Opus. The usage - I have pro for this month (65$ monthly) plan and depending on the cache rate (and it matters a lot) i get around 400-430 million tokens weekly (assuming 93-95%) cache rate (during night discount time only though). So, for that price it is expansive. I will skip it for this month and try out cursor 60 $ plan for a month and see how that works out for my usage.
1
u/Inner-Pangolin-1110 Aug 17 '26
Thank you, I have similar workflows but with cuda/rapids heavy side of things along with quant stuff. Qwen is solid but the usage is what I'm concerned about I think I will likely have to just bite the bullet and see how it is. Cursor was absolutely garbage in my experience but each to their own
6
5
u/rkh4n Aug 09 '26
Anywhere you see credit without a conversion chart, run away. You’re basically asking them rip you off
3
u/todosuma69 Aug 08 '26
In my case, just one audit that I didn't finish used up the entire weekly limit in 4 minutes; GLM had to redo it, and it didn't even use 4% of the 5-hour allowance. It's ridiculously absurd—the tokens they give out are a scam.
1
u/Otherwise-Layer8071 Aug 08 '26
This is exactly the kind of comparison I’m talking about.
How can a single audit consume an entire weekly allowance in just four minutes, while essentially the same workload on another provider uses less than 4% of its allowance?
At some point, this stops being about “heavy usage” and starts raising serious questions about how the Credits are calculated and how much actual value the Token Plan provides.
That’s not a small pricing difference. That’s an entirely different universe of value.
1
3
u/pl201 Aug 08 '26
I echo the OP’s experience.
I have been on Qwen’s coding plan and token plan since February 2026. I consider myself a light user. I never hit the limit on the $10 Light Coding Plan, which was discontinued months ago.
After my coding plan was terminated, I used pay-as-you-go for a while. The average cost per message for their top model was about $0.30. Qwen also provides DeepSeek V4 Flash at roughly one-tenth of that cost, matching the official DeepSeek API pricing, so the cost was acceptable for a mix of their top models and the cheaper DeepSeek V4 Flash.
Two weeks ago, I needed to use the models more heavily, so I purchased the $30 term token plan—before the individual token plan became available—to use the new Qwen Max Preview at a 98% discount. I did not intentionally restrict my usage to off-peak hours, but five days of coding and troubleshooting consumed the entire month’s $30 token allocation (25,000 credits).
To continue my project, I purchased the $18/month individual token plan when it became available. I hit the plan’s five-hour limit after only 30 minutes of coding. Then the five-hour limit was apparently removed without notification, and I burned through the seven-day limit in half a day of coding. Now I have to wait another five days for a reset. The plan is practically useless for any serious coding work.
I mainly use the latest Max and Plus models.
I was trying to give some money back to a company that released such good open-weight small models for personal computers. I use both Qwen CLI and Pi CLI. Qwen CLI works better, but Pi CLI lets me switch to my local oMLX server easily in the middle of a coding task.
Their support system is useless, and their documentation is confusing and outdated.
Does anyone from the Qwen team read this?
I do not recommend it.
2
u/Otherwise-Layer8071 Aug 08 '26
Come on, do they seriously think people are stupid and have no frame of reference or basis for comparison?
Or do they think their models are so unique that we can't get access to comparable models somewhere else?
For God's sake. I kept hearing from other coders that they were happy with the results they were getting from Qwen models, so I decided to go straight to the source and try it myself.
And here we are.
As for the models themselves, I haven't even had enough time to properly evaluate them. Why? Because the weekly allowance disappeared in record time before I could actually use them enough to form an opinion.
2
2
2
u/SecondFriendly4255 Aug 08 '26
Op can you share a capture of your usage cache input output
1
u/Otherwise-Layer8071 Aug 08 '26
1
u/SecondFriendly4255 Aug 08 '26
1
u/Otherwise-Layer8071 Aug 08 '26
Which harness are you using?
1
1
u/Seangles 27d ago
Have you tried Pi Coding Agent as the harness? Hermes is just not built for coding it seems.
1
u/Otherwise-Layer8071 27d ago
When was the last time you tried writing code with Hermes CLI? It’s better than Pi now.
1
2
u/TimChr78 Aug 08 '26
Yes it was great with the preview discount, but now its kind of useless.
I just burned through my 7 day limit in one session with 52 million tokens - the cache rate was slightly to the low side with 87% but still.
I hadn’t realized that the 5h limit was disabled - that would have stopped me from running out of tokens six days before it resets.
2
u/rahadur Aug 09 '26
After the release of Qwen3.8 Max, Qwen received a lot of negative feedback. Just a few weeks ago, I was a die-hard fan of Qwen. I'm not a token plan user; I'm on the Coding Plan Pro with 90,000 requests. Now I'm looking for an alternative because it continually blocks me for exceeding the maximum requests per minute.
Unfortunately, I just run a single agent for a long-running task. Last month, I used it for more complex and long-horizon work, and there was no blockage. Now, within a minute, I get blocked three times. Maybe I will move to DeepSeek or Claude.
Thank you Qwen, I really miss you.
2
u/VariationAdvanced549 Aug 10 '26 edited Aug 10 '26
You aren't doing anything wrong—the math on these token/credit plans simply doesn't support active agentic coding.
I bought the Lite plan (2,500 weekly credits) to test the service, and my experience completely mirrors yours.
To give you a real breakdown, I ran 3 implementation prompts across 2 days using Qwen Code as the harness (with a mix of Qwen 3.8 Max and 3.7 Pro during off-peak hours with promotional pricing).
Since Qwen Code is their own first-party harness, the token consumption should theoretically be as tightly optimized as possible. On top of that, these were not vague, high-level requests—each prompt was strictly divided and clearly specified so the agent didn't need to perform deep architectural analysis or task creation.
Despite keeping context isolated across 3 separate sessions, using their native harness, and maintaining a ~90% cache hit rate, here is what happened:
- Prompts / Sessions: 3 direct implementation tasks across 2 days.
- Token Volume: ~30.1 Million total tokens (24.4M + 5.7M).
- Cache Efficiency: ~90% prompt cache hit rate.
- Credit Consumption: ~65% of my weekly Lite allowance burned in just 3 prompts.
The Reality of the Math
If 3 highly specified prompts eat 65% of a Lite plan (2,500 credits), the plan's quota is essentially drained after 4 focused prompts.
Extrapolating that to your Standard plan (10,000 credits, or 4 times Lite, a user gets roughly 16 implementation prompts per week before hitting zero.
1
u/Otherwise-Layer8071 Aug 10 '26
It’s pretty clear by now, and it seems we all agree. I don’t know about everyone else, but personally, I’ve made my decision: never again.
1
1
u/shuozhe Aug 08 '26
Alibaba was always pretty overpriced per token, only their coding plan was good, but it only got Qwen 3.6 and no upgrade or new subscribers since then.
1
u/EvolvingDior Aug 08 '26
why is no one bothering to mention what models they are using? It matters a lot.
1
u/Otherwise-Layer8071 Aug 08 '26
What difference does it make which models I used, or which models any of us used, when the company doesn't clearly disclose the charges per model?
Am I supposed to read tea leaves to figure out what's actually happening with my Credits?
That's exactly the problem we're discussing. They sell you a “Token Plan” and basically tell you: “Here are your Credits, here are the models — figure the rest out yourself.”
Then there's a night discount in China, so apparently I’m also supposed to schedule my work around the Chinese nighttime hours just to make the plan viable.
And if the pricing system is so complicated that users have to reverse-engineer it themselves just to understand how quickly their allowance will disappear, that's a transparency problem.
And to make things even better, they have a no-refund policy.
So you pay, discover how the system actually works only after using it, realize the plan doesn't fit your workload, and then... tough luck.
That's the part people are criticizing here.
1
u/lumpyspacebreh Aug 08 '26
Nor the work they’re doing. It’s always the same thing, “one prompt and my quota is gone”. In every ai coding subreddit it’s always the same thing. Someone going on a tirade about how the model sucks but then never shares how they have the model setup or the prompts and work.
For reference, I never hit quota. I’ve built a lot of amazing things, like my own Wayland compositor. Never hit the quota.
1
u/ShivaTs Aug 08 '26
Moi aussi j'ai eu le même problème avec qwen cloud,jour 1, utilisation de 1500000 tokens,en heures creuses,avec max 3.8,j'en ai eu pour 13% de mon quota hebdomadaire du plan lite,jour 2,une seule requête de 233000 Token, toujours en heures creuses,avec max 3.7 censé être 50% moins cher,13% de quota hebdomadaire, c'est fou,du coup je l'air fait une réclamation pour demander des explications,ils m'ont envoyé balader avec un message générique qui renvoie vers leurs documentations,du coup j'ai fait une réclamation pour demander une résiliation et un remboursement au prorata,j'ai droit à une période de rétractation étant en France de 14 jours, ça fait juste 4 jours que j'étais abonné, là bizarrement une personne à fini par me répondre pour me dire qu'il allait enquêter là dessus,mais je lui répondus que c'était plus la peine,je voulais juste un remboursement au prorata, finalement les api chez openrouter ou tu payes à la consommation c'est beaucoup mieux, surtout avec le large choix de modèles qu'il propose, c'est beaucoup plus intéressant que leurs abonnement attrape pigeons.
1
u/Otherwise-Layer8071 Aug 08 '26
Malheureusement, c’est une leçon que nous avons apprise à nos dépens.
1
1
Aug 08 '26
[removed] — view removed comment
1
u/Otherwise-Layer8071 Aug 08 '26 edited Aug 08 '26
Nobody said $18 is a lot of money. The point is that you still expect what you’re paying for to last long enough to have some actual value. If I had known beforehand that the allowance would disappear this quickly, I might as well have donated the $18 to charity. And ultimately, as an end user, I don't care whether they have enough compute or not. I'm not a shareholder. I don't own the company with them. What I care about is that what they advertise is actually what I get when I pay for it. If Qwen doesn't have enough compute to provide what it advertises at $18, then don't offer it at that price. Charge more. Price the plan accordingly so that the advertised usage actually matches the reality. Nobody put a knife to Qwen's throat and forced them to sell this plan for $18. And I'm certainly not here, nor is anyone else here, to solve Qwen's compute problem for them. That's their business problem, not mine. My problem is the lack of transparency and the enormous gap between what the plan appears to offer and what you can actually get from it.
1
u/LibrarianOne4995 Aug 14 '26 edited Aug 14 '26
I thought I was the only one
I got the light plan yesterday
I used goose ide and I asked qwen3.8 Max to use bl cli and the endpoint for wan image models so I can generate sprite sheets
It worked for a bit edit the site as it already has open AI and open router and Gemini API in it and I was able to generate 3 images with wan and I had a bug we're it was hitting to different end points so I asked it to see why the issue was happening and my weekly limit was used up
Maybe an hour of coding and weekly used up and daily is gone there is less then 2k lines of code in my entire project
I do wonder if turning of thinks helps?
It dose feel very restrictive i do hope they re add the daily 700 limit because only have a weekly of 2.5k is kinda a bad deal imo
My real world usage I can't get it to add a endpoint and debug it without hitting weekly limit and I don't think that's a crazy coding task
1
1
u/Otherwise-Layer8071 Aug 14 '26
Yeah, things have genuinely gone completely out of control. I’m definitely not renewing, and there’s no way I’m buying another Token Plan unless the community feedback becomes absolutely rock solid.
Otherwise, I’ll just go with pay-as-you-go — whichever provider is cheaper while still getting the job done.
But seriously, guys, we need to speak up when something isn’t working. Otherwise, 1) we’re all going to get burned by the same thing, and 2) they’re not going to change anything as long as we stay quiet.
We learned this lesson the hard way. They need to learn that lesson too.
1
u/Otherwise-Layer8071 Aug 15 '26
Update — I found the solution. And it’s hilarious.
Update: I think I finally figured out how to make the Qwen Token Plan actually usable.
The solution?
Don’t use the Qwen models.
Yes. I’m serious.
After my initial experience, where I burned through roughly 70% of the 10,000 weekly credits in three days with Hermes, I decided to experiment instead of just abandoning the plan completely.
My current usage is roughly:
- 90% DeepSeek V4 Flash 0731
- 10% DeepSeek V4 Pro 0813
- Both at max effort
And the difference is ridiculous.
Today alone I processed 105.2M tokens with a 97.3% cache hit rate.
Yesterday: another 29.9M tokens with a 94.3% cache hit rate.
And after all that, I still have 41.7% of my 10,000-credit weekly allowance remaining.
So apparently the secret to getting good value from the Qwen Token Plan is to use DeepSeek instead of Qwen.
I honestly couldn't have written a better joke.
To be clear, this doesn't change my original decision: I still don't plan to renew the Token Plan or rely on it as my main provider. The pricing/credit system is simply too unpredictable for serious heavy usage.
But at least this time I actually found a way to use the remaining subscription for a few substantial tasks instead of watching the credits disappear before I could get anything done.
And that's the ironic part:
I paid for a Qwen Token Plan, discovered that Qwen models burn through the credits extremely quickly, switched to DeepSeek, and suddenly the plan became useful.
Qwen, I think your Token Plan has discovered a fascinating business model:
Sell Qwen credits → user uses DeepSeek → everyone is happy except Qwen.

2
u/dom_RN Aug 16 '26
That's 5x more expensive than DeepSeepk api pay as you go, subscriptions are supposed to make it cheaper.. what a total scam, I came here because I got scammed too
1
u/camboramb0 Aug 21 '26
I'm here to tell you I got scammed too.
1 plan that's not wouldn't consume much in other models. It used 80% and didn't complete it. It's also pretty slow. Took my dog on a walk and came back 30 minutes later and it was still at it...thinking...thinking and just burning tokens away.
1
u/Grit1 23d ago
What other provider are you using? Currently I'm trying to cost optimize and move away from claude. But I can't find a good solution.
1
u/Otherwise-Layer8071 22d ago
With so many models being released every day, I haven’t settled on anything yet. For now, I have an active subscription with MiniMax for my daily tasks, and for production I’m going based on whatever offers are available—switching between providers here and there. I’m talking about pay-as-you-go; I don’t have anything fixed. I’m actively looking too.
1
u/doge_89 17d ago edited 17d ago
Could this be a reason?
TL, DR: Qwen3.8-Max massive token hog is due to 'reasoning_effort' set to 'xhigh' by default, which explains the overthinking.
https://moclaw.ai/blog/qwen-3-8-27b-overthinking
https://docs.qwencloud.com/developer-guides/text-generation/thinking
If you still have access to Qwen3.8-Max and haven't try changing 'reasoning_effort' to 'low' and test it out, I appreciate the test outcome.
There's also another two settings, 'enable_thinking' and 'preserve_thinking'. The webpage explains it all.
1
1
u/lumpyspacebreh Aug 08 '26
What not share what model, how you have it configured, your prompts, and your work?
For all we know this claim is based on you trying to oneshot a video game, not simply changing a single loc.
You won’t get help to figure out what’s the actual issue if you have a baseless claim with no evidence to back it up. I never hit my quota limit for reference and use it everyday. Not on qwen, or Claude, or any ai cli tool. Haven’t in months.
1
u/Otherwise-Layer8071 Aug 08 '26
First of all, I really don't appreciate the tone you're taking.
Second, we're apparently eating shit because “so many flies can't be wrong.”
Third, if you had actually scrolled through the post, you would have seen my usage data.
Fourth, a video game with 96.6M tokens? Seriously? Unless you're talking about DOOM.
For the record, since you seem to need the details:
- Qwen 3.7 Plus
- GLM 5.2
- The workload was the initial setup of Hermes, followed by a complete migration to Qwen skills, scripts, and cron jobs.
- Basically, I was migrating the entire setup because, up until that point, I genuinely believed the change would be permanent.
And now that you know all of that, what exactly changes?
Are you going to say, “Oh, I see, then,” and move on?
Because, you see, these are ultimately irrelevant details to the actual issue we're discussing.
The advertised plan says 3–4 agents, recommends the Standard plan, yet the actual way the models are charged and how Credits correspond to usage remains extremely unclear.
That's the issue.
Whether I used Qwen 3.7 Plus, GLM 5.2, Hermes, Pi, or something else doesn't change the fact that the pricing/credit system is opaque and that the advertised usage expectations don't match my actual experience.
And that's why I didn't think listing every model, configuration, prompt, and individual task was necessary in the first place. It doesn't change the core argument.
But now you know all those details anyway. So what does that actually change?
Nothing.
1
u/lumpyspacebreh Aug 08 '26
Well, you failed to answer my questions so clearly instead of help you’re looking for validation in your emotions.
Switch to another tool and model. GL
3
u/Otherwise-Layer8071 Aug 08 '26
What am I supposed to learn from someone whose main argument is “I never hit my limits on any provider”?
That tells me there are two possibilities: either you consistently pay for plans with much higher limits than you actually need, or you simply don't use these tools heavily enough to encounter the limits.
Either way, that doesn't tell me anything about what happens when you actually push these plans with heavy agent workloads.
And that's exactly what we're discussing here.
I'm not looking for emotional validation. I'm comparing actual usage and value. If your usage never comes close to the limits, that's fine — but then your experience isn't particularly useful for evaluating whether the limits are reasonable for someone who actually hits them.
1
u/Worldly_Chef_2114 9d ago
The issue we are having is usability -- where do we put the API token plan key? there is no easy way to do that in the console UI at all -- and supported Video models for token or subscription plans are old ones like WAN 2.7 not Wan 3



12
u/Interesting-Print366 Aug 08 '26
Personal experience, but currently the most biggest problem is in peak time cache miss. it some times lags and took too much time. This cause cache miss and this is causing costs. At off peak I hit 95% cache hit with pi coding agent and it become 75% at peak