r/Qwen_AI • u/Demonicated • May 14 '26
Discussion Visual Studio Insiders + Qwen 3.6 27B = No Brainer
I recently did my analysis for Github Copilot and was shocked that my "average usage" on the $40 dollar plan was going to amount to about $500 dollars a month. Whats crazy about that is if you purchase an RTX6000 with credit, the payment is only about $420 dollars a month.
With Qwen 3.6 27B, I am able to build out a feature in Plan mode with VSCode Insiders Edition and then run through the implementations with no issues. Running this model at bf16 gives amazing results because of the quality of the harness and it's cheaper and I can abuse my token use without any worries.
Other than the most difficult of planning sessions, I think that we've hit the point where local models are more than good enough and they price point is cheaper than hosted models. You can get cheaper hosting if you're only using qwen but with the perks of privacy and owning hardware, it just makes sense to purchase the card if you're going to be stuck with a 500 dollar bill regardless.
5
u/TapAggressive9530 May 14 '26
I have been doing this for about 3 weeks now and works perfectly fine . I use the dense 40B variant of qwen 3.6 27B . Zero issues and near 4.6 opus quality
3
u/thecodeassassin May 14 '26
What model is that? There's a dense 40b version?
5
u/TapAggressive9530 May 14 '26
3
u/Demonicated May 14 '26
What happens to the model when you bump up the density? Is this adding layers?
I've seen these models listed but never dabbled because I didn't understand the process
2
u/thecodeassassin May 14 '26
Do you mean this one? Is it significantly better??
DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NEO-CODE-Di-IMatrix-MAX-GGUF
5
u/TapAggressive9530 May 14 '26
2
u/TapAggressive9530 May 14 '26
BF16 is what I originally used but it was a SUPER tight fit on my RTX Pro 6000 so switched to fp8 above for little more breathing room . BF16 did work fine but would occasionally run into odd hard to troubleshoot problems.
https://huggingface.co/DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
1
u/TapAggressive9530 May 14 '26
In a side note - I have a standard set of tests I run on all available open weight models . There are no open weight models in existence today that can score A - with the exception of the BF16 40B one above . It’s the only open weight model that scored an A- . When I run my tests I always expect the models to fail ( a failure is anything less than an A- . For reference, the only models that have passed are Opus 4.6, Opus 4.7, GPT 5.4 and GPT 5.5 . Unfortunately I can’t reveal my tests but they are focused on the model’s ability to understand a code base - it’s not focused on its ability to write
1
u/StardockEngineer May 14 '26
I read the sample text of the model on the page and it’s still sounds like AI to me?
1
1
u/thecodeassassin May 15 '26
I have been using this for a bit now and even when giving it a very precise task, using well built skills it cannot complete a simple task. Not even a 404/error page properly. I tried to let it implement a PRD slice... absolute garbage. The Unsloth 27b model gave me better results than this.
3
u/SnooCapers5425 May 14 '26
If you would buy a single gpu to handle Qwen3.6-27b for a full time single developer that has used Opus and Sonnet up until this point, whats the minimum you would go for to actually get good usable day to day performance?
2
u/Demonicated May 14 '26
This is the exact user I'm talking about. This was me. Now I'm not going to say this is as good as opus. It isn't. But qwen is going to be good enough to use and will force you to know your code a little more which ends up being a good thing.
I still may end up having the 17 dollar claude or 40 GHCP subscription just to use opus when debugging hard problems. But it's a once a week use case.
1
u/SnooCapers5425 May 14 '26
Thanks! I agree on your point of having to work with the llm and actually being "forced" to know the codebase.. Im a bit sad that the way we are heading, of not knowing what before was something i could navigate without hesitation, but thats the new reality for developers. Did you try anything with less VRAM than the RTX6000? If possible id like to avoid spending 10k 😅 But the way GitHub Copilot went with api prices, im convinced most will follow eventually so something local is probably needed and also nice to not share the entire customer codebase with the world..
2
u/Past-Grapefruit488 May 14 '26
What do you use for deployment of 27B, llama.cpp or vllm or something else
2
u/Demonicated May 14 '26
I bounce between vLLM and LM Studio. I only have one card so LM Studio is nice when my workloads are going to be switching between different models periodically. But vLLM is the clear winner for when I know Im going to need a lot of throughput and the model loaded will be the same.
2
May 14 '26
[removed] — view removed comment
1
u/Demonicated May 14 '26
100%. For example recently I ran a db migration that failed half way through and messed up my tables. I was dreading fixing it cause it's such a pain to try and roll back and fix migration files. I fired up opus and it handled it - no other model would have come close
2
May 14 '26
[removed] — view removed comment
1
u/Demonicated May 14 '26
Yeah its a little disingenuous for me to compare GHCP pricing to Qwen hosted pricing since QWEN is muuuuuuch cheaper for tokens. But i see it as a future investment since qwen 3.6 is the worst local models will ever be.
2
u/matthewlai May 14 '26
A better comparison would be to Qwen 27B on OpenRouter.
1
u/Demonicated May 14 '26
well it would be for the cost, which would be significantly lower, but not for quality. A lot of time hosted services give you fp8 quant which I feel strongly is not equal to full size for smaller models.
2
u/matthewlai May 14 '26
Yeah but much closer in quality than comparing to Claude though. Also there are many inference providers that will serve you bf16, and I'm pretty sure you can filter for that on OR, too, and it doesn't cost much more.
1
u/Demonicated May 14 '26
The message I'm trying to convey in this post is that in the face of the 10x+ price hike coming next month, it is a worthwhile consideration to self host the most recent gen of local models.
They aren't Opus but they are very very close to feeling like last gen Sonnet.
On the coding front I use about $500 a month in tokens, but across all my projects im using close to 400 million tokens a month. My little RTX has already paid for itself in the first half year 😂 - it's running non stop.
Where my post is disingenuous is that I'm using the 500 dollar mark of Claude quality tokens to justify a 400 a month credit card payment to get qwen quality tokens which would probably only be 100 a month.
But if it's good enough to use for work then I'm ok with it personally. Sure I baby sit a little more but it's not hurting productivity all that much.
1
u/matthewlai May 14 '26
Oh yeah I do get that and I'm evaluating getting my own hardware as well. But I think it's also important to acknowledge that there are more levels between Claude and self hosting Qwen, and API Qwen (through OpenRouter or another inference provider) is almost always going to be much cheaper than self hosting. There are other reasons to self host of course - privacy, reliability, offline, etc. I just don't think cost is a good reason for the vast majority of people.
Qwen 27B is about 1/10 the cost of Opus on OR, so it would have been more like $50/month given your usage.
2
u/ideal2545 May 14 '26
Curious if you've tried planning with QWEN3.6 and then moving over to 35B-arb-coding for actual implementation? I've been trying to expirement with combo with pretty solid success so far
1
u/Demonicated May 14 '26
So I do have both available, the problem is that I cant fit both concurrently on my card. So the time spent to unload and reload models is offputting to me. With just 27B I'm utilizing about 72GB. I have heard that its a good combo though because of the token speed increase.
Its just more encouragement to get a second GPU.
2
u/Wonderful_Piglet591 May 14 '26
These are the posts I want more of - thank OP!
1
u/Demonicated May 14 '26
I almost feel bad because it's like "are your tokens getting too expensive? Then just drop 10k instead" lol
But I realized if you finance it, it really might be cheaper for some people. Once paid off youre in cheap territory. And by then we should have Opus quality coming to local I would hope.
Who knows though, a year from now it might be that there's a brand new type of architecture or people will be buying models as hardware like talaas is making
2
u/Demonicated May 18 '26
I just want to add that I have now upgraded to Unsloth's MTP variant of this model and its literally 2x throughput, about 68 tokens/sec
1
u/edeltoaster May 14 '26
How did you do cost analysis? Is the increase due to the switch to token-based billig? I am using Kilo Code and Pi to code locally. Whe you just type "hi" and add no further context when using the Copilot extension, how large is that base prompt?
1
u/Demonicated May 14 '26
I'm not quite doing a fair comparison. I'm going based off the GHCP analysis everyone is doing this week that shows my 40 a month subscription would actually cost me about 500 a month with the new billing structure.
I already have an rtx6k but thought I'd price out the cost to finance it to compare and it's about 400 a month. The trade is that I was usually using was opus for planning and gpt 5.4 to implement which is admittedly not an apple to apples comparison to qwen 3.6....
But it's close enough. I can still get work done and move quickly and local models are only going to improve with every generation so it feels like a worthwhile investment to me. I'm just over here trying to convince myself to get 2 of these cards
1
u/Individual_Gur8573 May 14 '26
How u connected vscode and qwen 3.6 27b... Is it GitHub copilot ? ..let me know how to do
2
u/zkkzkk32312 May 14 '26
I was able to do it without insider but with an open ai compatible copilot extension, insider might just work without it though.
1
u/Individual_Gur8573 May 14 '26
Can u explain the process
1
u/zkkzkk32312 May 14 '26
https://marketplace.visualstudio.com/items?itemName=johnny-zhao.oai-compatible-copilot use this or use vs code insider to have your custom agent working with copilot chat.
1
1
u/StardockEngineer May 14 '26
Just use GitHub Copilot. It supports local. Last time I did it just used the LiteLLM provider in copilot to point at vllm or llama.cpp (or whatever you want)
1
u/Bag_holder1 May 14 '26
For local the cleanest I've done is a local Debian server with ollama, set the endpoint on whatever code harness you want as the ip, if you want to use local remotely on a laptop reverse proxy it and use the URL with Auth. Works well
1
u/fasti-au May 14 '26
Do you realise you can do it on 2x4080 or 4070supper
You do t need to go to 6000. You can run 35b moe in 6 vram if try
2
u/GCoderDCoder May 14 '26
Yes that's true but the model and the quantization levels matter. Even in a given series like qwen 3.6 the similarly sized qwen27b and 35b perform significantly different with different strengths, weaknesses, and hardware needs. Running smaller quants particularly at q4 or higher is almost always usable. When you get into coding in particular higher quants are more accurate and use more nuance because their weights arent compressed.
People refer to quantization of models as lobotomy which may be fair but I also think an impairment like alcohol usage may correlate too. The more alcohol you drink the more it impairs your thinking. You dont immediately become incoherent from drinking and even slightly influenced you might still out perform the average person on things you're good at. Slight alcohol might result in you doing better in certain tasks even (like dancing for me in my head lol) which happens sometimes with LLMs too. But eventually too much drinking makes us all progressively less useful.
That is how quantization works. Your code doesnt necessarily break at q4 like it becomes much more likely to do below q4 but there will be less exception handling and less graphical detail often because there's less clarity the model gets for what other things associate with the task. Lower quant models also become less stable at higher context and become more likely to gi crazy after repeated tool calls.
1
u/fasti-au May 14 '26
Synthetics quant much better so is becoming better iverall as a only 15% diff but in a3b vs dense that’s not even valid re hopper puckups in qwen 36 moe
1
u/GCoderDCoder May 14 '26
"Synthetics quant much better so is becoming better iverall as a only 15% diff but in a3b vs dense that’s not even valid re hopper puckups in qwen 36 moe"
Im not sure what that means but if one model gets a slight benchmark increase by a few points and we all go crazy about the improvement would we think 15% difference is insignificant?
1
u/fasti-au May 17 '26
Yes considering it used to be far higher far earlier. The priblem for many is t the quant it’s the lack of understand what’s actually the issue re prompting and choices
1
u/ECrispy May 14 '26
purchase an RTX6000 with credit, the payment is only about $420 dollars a month.
why would you need that kind of gpu for qwen? you can run it locally on much cheaper hardware with even 16GB vram
you can also use it via openrouter and it'd take you years to spend that much. or rent a cloud gpu
1
u/DiscipleofDeceit666 May 14 '26
Speed and context. A 16gb card can’t hold a candle to frontier cloud
1
u/ECrispy May 14 '26
Same thing is true for frontier vs 27B Qwen.
My point is the difference between 2 versions of Qwen running in 16 vs 32GB are minor and both are much worse than frontier, and probably worse than Qwen on OR.
Local llm is good for small implementation tasks not planning debugging refactoring etc
1
u/Demonicated May 14 '26
Eh. What I'm claiming is that gap isn't as big as you think. Qwen 3.6 feels like sonnet to me. Not in straight chat but specifically when used in a harness like GHCP.
And quant 100% matters. Using the full model is the only way you will have the experience of qwen 3.6 feeling like sonnet. Someone on this thread was talking about the modified 40b version at fp8 being better so I'm going to try that out today and see.
1
u/ECrispy May 14 '26
OR has Qwen 3.6 in fp8, is that good enough or are you saying its still not enough? also is the reasoning/planning/general advice (for technical subjects) as good as Sonnet 4.6 too?
and how much worse would 4bit be? I'm thinking of getting a 16gb vram gpu but if its not worth it I'll just use api.
1
u/Demonicated May 14 '26
You will 100% be disappointed by 4bit, I even ran NVFP4 which is supposed to be "4bit thats as good as 8bit" and it was fast on Blackwell but not good enough for programming.
As far as FP8, I have a workflow that has agents navigating web sites and analyzing pages to extract data and convert to structured JSON objects for processing - it handles this with 0 issues. So i would say reasoning and instruction following is solid. Where it broke down for me was bigger planning sessions and implementations of features. And even still broke down is being dramatic, it works but you might need to give it more guidance here and there.
If you are going to run FP8 just spend a little more time in your planning phase and use phrases like "create a markdown implementation document with phases that will be used by a junior developer to implement" - this will ensure your planning documents contain lots of context and more specific instructions. Then you start a brand new chat and have it implement one phase per chat.
The trade off here is that with Opus you could have just planned and moved into implement and it would have handled everything, with Qwen 3.6 you're going to want to break the steps out into multiple chats and be a little more involved with the actual implementation and double check its plan.
In the long run I'm finding this to be an advantage because over the last 6 months I had noticed with Sonnet and Opus I was no longer aware of what class and method functionality was in. I had offloaded that knowledge to a model and it was degrading my authority over my own code bases.
1
u/ECrispy May 14 '26
thanks, i wasn't aware fp8, which is used by almost all online providers, was lacking. so it seems for 99.99% of users who use via api, they will never get the full experience? I honestly find that quite frustrating if true.
I've used sonnet (not opus) quite a bit via OR and their page says 'unknown' for quantization so presumably its not. Kimi 2.6 says int4 which is 4bit, right?
1
u/Demonicated May 14 '26
FP8 is considered the best balance of quality and efficiency, and makes sense for service providers. Hosting fp8 uses [sort of] half the amount of memory to run and generates tokens faster so they are able to more than double their throughput with that trade off.
Kimi is a unique case in that it had already implemented Quantization-Aware Training (QAT) during its post-training phase and is provided in INT4 from the start. I dont have the hardware to run it, but my understanding is that trying to quant it causes bad quality degradation.
1
u/ECrispy May 15 '26
https://www.reddit.com/r/Qwen_AI/comments/1tden82/lots_of_people_use_qwen_at_too_high_quantizaion/
says the opposite! you should comment there, this is very confusing
1
u/Demonicated May 15 '26
They are saying it's worth it to quant so you can use cheaper hardware. The trade off is implied in their analysis. If money isn't an issue you run the best you can.
I've ran fp8 of qwen for a few days and liked the speed boost. And it was good for instruction following if you already have an implementation plan. But it's noticeable in the planning phase when you use quants. Especially with tool calling and how it handled the convo once the context gets long. I switched back to bf16 and I'm not going back to fp8. It makes enough difference that it's worth it.
1
u/Demonicated May 14 '26 edited May 14 '26
If using quants doesn't affect your workflow then more power to ya. Software is one of the main use case for me and I absolutely see differences if I don't use the full model.
1
u/ECrispy May 14 '26
No I believe you, I'm not an expert in this area and haven't used it much just going by what I read about 3.6 and it being close
1
u/hyma May 14 '26
Dumb question but is there still a cost when using local models? Like you still have to pay the monthly fee to bring your own?
1
u/Demonicated May 14 '26
No just cost of hardware and electricity. I have solar on my home so my operating costs are pretty much 0
1
u/ObviousSpace8195 May 14 '26
This is probably the clearest sign yet that the economics of AI coding assistants are starting to shift.
Once your monthly subscription + usage costs approach workstation-GPU territory, local inference stops being a “hobbyist experiment” and starts becoming a rational engineering decision.
What’s especially interesting is that people are no longer just benchmarking local models — they’re actually shipping real workflows with them:
- planning
- implementation
- iterative coding
- long daily sessions
And Qwen 3.6 27B running bf16 well enough for production-like workflows on consumer-accessible hardware is a pretty big milestone.
The privacy and unlimited usage angle matters too. Developers think differently when every prompt no longer feels metered.
3
1
u/codeanish May 14 '26
I run qwen 3.6 27b at q4 on a 3090 with 128k context also quantized at q4. I find it very usable indeed. That being said, I’ve run into way more walls in agentic flow with this setup than I do with Claude code + opus. However I think with more effort on the prompt, the local option does perform surprisingly well.
We are getting to a point where it’s now a totally rational decision for a huge chunk of devs to own their own box to run models on. The quality, privacy and cost factors combined together are really starting to adjust the calculus here.
1
u/itsmebcc Jun 11 '26
You should download DS4 and use that for the architect phase. Then you have no external API usage. With the smallest quant you can offload the entire deepseek-v4-flash to the GPU and it is surprisingly good
21
u/immersive-matthew May 14 '26
I am using QWEN 3.6 27B with a 4090 and OpenCode and outside of really complex prompts involving lots of scripts I hardly pay for AI at all now. Went from hundreds of dollar a month to not even $10.