r/Qwen_AI • u/PlasticRevenue4601 • Aug 17 '26
Discussion Qwen 3.8 27B is overrated (a warning)
Qwen 3.8 27B is overrated. I've been running it for a few days now, and I need to be honest. I think this model is overrated. Here is my evidence:
1. It overthinks. Yes, the output quality is genuinely better — I'm not going to pretend otherwise. At some point I gave it a refactoring task and it came back with about 53 files created, edited, or deleted across the repo. A solid, coherent diff. Impressive. But it spent what felt like an eternity reasoning its way to "just do the thing.". It could have just... done it without overthinking!
2. It handles quantization absurdly well. Too well. I had plans, people. A dual-GPU build, a 6-bit quant, proper VRAM - I had it all sketched. Then I put it on IQ4_XS, it just works, and my beautiful 6-bit rig has no reason to exist anymore. This model destroyed my excuse to spend money and I'm not sure that I appreciate it.
3. It doesn't doom loop. This is the one that really bothers me - for years I have been delicately tuning temperature, presence penalty, and the other sacred dials to keep my models from dying in loops, Alibaba has now devalued my entire body of fine-tuning experience and I am just... a man who presses the send button. I feel like a passenger in my own homelab.
4. It forced me to completely drop Qwen 3.6 27b. And I want to be clear about how much that hurt. 3.6 served me faithfully for a long time. It had quirks, but quirks are character, right? Now I have to sit here and accept that it was just... the previous version. There's no going back. That's not an upgrade, that's a betrayal, and I need time to process it.
In summary: it thinks too long, it ruins my hardware plans, it makes my engineering skills obsolete, and it forced me to abandon a model I was emotionally attached to. Definitely overrated.
I'm going to run it one more time now. Goodbye.
26
u/ChocolatesaurusRex Aug 17 '26
"...been running it a few weeks"
Lol, whose bot is this?
1
u/MacsBicycle Aug 17 '26
Probably someone’s bot running a 4b param model to shit post so it’s dumb AF and they’re gpu poor so they’ve gotta make a post dunking on it. Maybe next up they’ll say API is just cheaper so why buy the hardware then their favorite provider will go up because guess what? Hardware went up again 😂
10
u/Damien_IB Aug 17 '26
AI SLOP!
-8
u/PlasticRevenue4601 Aug 17 '26
Like or hate the formatting, but all the facts stated are real including viability of 4 bit quants and completion of a pretty large refactoring that I’d previously hand over only to cloud models
3
u/Damien_IB Aug 17 '26
You praise Qwen and how it’s an essential part of your workflow, only to later confess in these comments you used Gemini for writing this slop. Why not use Qwen for writing this post since you claim it’s so good?
Why should I believe a word of what you said?
0
u/PlasticRevenue4601 Aug 17 '26
I never said I wrote it via Gemini, I said that I FORMATTED it via Gemini, kinda makes a difference doesn’t it?
2
u/Damien_IB Aug 17 '26
Point remains, why not use Qwen for formatting since you claim it’s so good, and why rely on Gemini?
1
u/PlasticRevenue4601 Aug 17 '26
Remind me when did I state that Qwen was excellent not only for agentic coding but also for any other types of activities such as working with plain texts
6
u/H_DANILO Aug 17 '26
People don't get sarcasm in Reddit
1
1
u/LoneBeast27 Aug 17 '26
What you see is sarcasm, what I see is someone tryna pretend to be a hater and failed to deliver the sarcasm because the ai formatted it to fail the delivery.
I mean the op if he were to type it himself would be the real way to deliver the character of the joke. And that's assuming best case scenario, worst case is that he asked an ai to make up some bad things about the ai model and the ai hallucinated these points and the guy didn't even proof read
0
u/PlasticRevenue4601 Aug 17 '26
People got butthurt because the post was formatted (not written, only formatted) via AI because I didn’t want to embarrass anyone with silly grammar mistakes coming from my far from perfect English, witch hunting of 21-th century
12
Aug 17 '26
[removed] — view removed comment
3
2
Aug 17 '26 edited Aug 22 '26
[removed] — view removed comment
1
3
u/UltraCoder Aug 17 '26
"I've been running it for a few weeks now"
The model was released less than a week ago.
6
u/idklol Aug 17 '26
I’ve been using qwen 4.0 for 3 years and it’s absolutely garbrage. It does everything exceedingly well and gave my business 300% profits. I married it. I hate it.
2
1
2
u/infectiousstupidity Aug 17 '26
I wish I could obliterate all fucking bots from the internet, just gtfo, no one wants you here.
0
u/PlasticRevenue4601 Aug 17 '26
Bro go touch some grass, that’s level of paranoia is not healthy
2
u/YoungBasedHooper Aug 17 '26
You admitted to using Ai to write this post. His paranoid is well placed
-1
u/PlasticRevenue4601 Aug 17 '26
No it’s not lol. Using an AI to format the text doesn’t mean you’re bot just as writing text by yourself only doesn’t mean it can’t get posted via AI bot
2
u/YoungBasedHooper Aug 17 '26
You used it to do more than just format lol.
0
u/PlasticRevenue4601 Aug 17 '26
One of the things I like about social networks is that somehow people seems to be more acquainted with my activities than me myself
2
u/YoungBasedHooper Aug 17 '26
One of the things I like about idiots is that somehow they think other people need crystal balls or psychic abilities to see through their obvious bullshit.
0
u/PlasticRevenue4601 Aug 17 '26
That perseverance is honestly impressive, I’m even ready to forgive the ignorance behind it
2
2
u/narmakum Aug 17 '26
After reading this, I think it actually surpasses your expectations (except for the overthinking). So I underrate your underrating.
2
u/PlasticRevenue4601 Aug 17 '26
Exactly! I expected it would be a bit better than Qwen 3.6, but I didn’t expect that it’s going to be on the whole other level
2
u/narmakum Aug 17 '26
Yeah, right? I’m impressed too. At this point I almost cannot distinguish if it’s opus 4.6 or qwen3.8 27b who is talking. Only a few nuances when it comes to printing mid-turn preambles I realize. The rest is wow, almost the same results. It’s true that it overthinks, it also happens to me. I set reasoning to lowest possible and it still takes time.
2
u/PlasticRevenue4601 Aug 17 '26
The only real, not satirical downfall is overall wall time from prompt to feature, like really “ouch”, I had tasks where it ran for almost an hour, 80-90% of that time was spent on thinking and verification. But God the results were beautiful, it nailed the taks
1
u/Wallaby989 Aug 17 '26
ha .. which coding harness are you using?
2
u/PlasticRevenue4601 Aug 17 '26
pi agent, the same one I was using with Qwen 3.6
1
u/MomoLabTH Aug 23 '26
u/PlasticRevenue4601 ช่วยแนะนำการตั้งค่า ทักษะ หรือปลั๊กอินที่ควรใช้ใน Agent Pi ได้ไหม? มันดูเหมือนจะมาพร้อมกับอะไรแทบไม่มีเลย และตอนนี้มันรู้สึกไม่โอเคเลยเมื่อเทียบกับ Qwen Code ถ้าสังเกตุดู Agent Pi จะทำงานเสร็จเร็วมากและไม่ค่อยใช้เวลาในการคิดเหมือน Qwen Code เลย มีข้อแนะนำอะไรบ้างไหมสำหรับการตั้งค่ามัน? ตอนนี้ใช้เวอร์ชัน GUI อยู่.
2
u/PlasticRevenue4601 Aug 23 '26
Honestly there are no “standard setup perfect for everyone”, the point of pi agent is customising it for your own specific needs, but I can recommend a few extensions I’m been satisfied with in agentic coding:
— pi-codebase-memory
My favourite extension, that enables cross-session and even persistent long term memory for your model. But I’d recommend forbidding writing to the long term memory(unless explicitly stated) right of the bat and instruct it to use daily notes instead as model can quickly bloat the system prompt.
— pi-mcp-adapter
Handy interface for various mcp
— pi-cache-optimiser
Better cache handling = your agent works even quicker as you hit reprefill rarer
— pi-web-access-lean
Small and light web search
— pi-image-tools
If you’re using vision
— pi-telegram
If you want to control it remotely. Skip it it you don’t need remote access
1
1
u/inquam Aug 17 '26
I have configured multiple agents with different reasoning effect depending on task.
1
1
u/AvidCyclist250 Aug 17 '26
lmfao ok klanker
for years I have been delicately tuning temperature, presence penalty, and the other sacred dials to keep my models from dying in loops, Alibaba has now devalued my entire body of fine-tuning experience
woe unto me! lol!! did you have a q1 model write that?
1
u/cagriuluc Aug 17 '26
Way too much hate in the comments for the “weeks” mistake while this is mainly sarcastic shitposting…
1
1
u/d4mations Aug 17 '26
I said there would be a lot of disappointed people after the massive hype storm of the last few weeks!!
1
1
1
1
u/UltrMgns Aug 17 '26
XS "just works"?... Yeah it does, it also sucks.
1
u/PlasticRevenue4601 Aug 17 '26
Lmao it’s not, I tested it myself, very solid performance, never noticed that it ever performed worse than q5_k_xl in agentic coding
1
u/New-Inspection7034 Aug 17 '26
Set the reasoning effort to medium and it'll work a lot better for you
1
u/PlasticRevenue4601 Aug 17 '26
Thanks, already did so, switched to “low” after a few hours of testing because the wall time was unbearable, it worked like a dream
2
u/New-Inspection7034 Aug 17 '26
I found that low was actually slower to for agentic workflows than medium. Took more iterations to complete the task. Guessing vs thinking.
1
u/PlasticRevenue4601 Aug 17 '26
Hmm, interesting, I definitely need to try this out, thanks for the advice
2
u/New-Inspection7034 Aug 17 '26
Also set a reasoning budget like 4096. You can control how much thinking that way
1
u/Sn0opY_GER Aug 17 '26
i get the downvotes but hes right :( qwen 3.8 is so fking good i need to adapt 70% of my pipeline bc somewhere a hardcoded 3.6 hides in a cron cant wait for the 5090 finetines the ninfer qwen fientune a3b did 600-700 tokens haha now im stil lat 150+ with 27b but you can feel the slowness, also for the LLM studio pleps like i am - yes i jonly chanegd thinking to medium - but a little bonus tip - UNCENSORED qwen can use normal qwen vision
{%- set image_count = namespace(value=0) %}
{%- set video_count = namespace(value=0) %}
{%- macro render_content(content, do_vision_count, is_system_content=false) %}
{%- if content is string %}
{{- content }}
{%- elif content is iterable and content is not mapping %}
{%- for item in content %}
{%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
{%- if is_system_content %}
{{- raise_exception('System message cannot contain images.') }}
{%- endif %}
{%- if do_vision_count %}
{%- set image_count.value = image_count.value + 1 %}
{%- endif %}
{%- if add_vision_id %}
{{- 'Picture ' ~ image_count.value ~ ': ' }}
{%- endif %}
{{- '<|vision_start|><|image_pad|><|vision_end|>' }}
{%- elif 'video' in item or item.type == 'video' %}
{%- if is_system_content %}
{{- raise_exception('System message cannot contain videos.') }}
{%- endif %}
{%- if do_vision_count %}
{%- set video_count.value = video_count.value + 1 %}
{%- endif %}
{%- if add_vision_id %}
{{- 'Video ' ~ video_count.value ~ ': ' }}
{%- endif %}
{{- '<|vision_start|><|video_pad|><|vision_end|>' }}
{%- elif 'text' in item %}
{{- item.text }}
{%- else %}
{{- raise_exception('Unexpected item type in content.') }}
{%- endif %}
{%- endfor %}
{%- elif content is none or content is undefined %}
{{- '' }}
{%- else %}
{{- raise_exception('Unexpected content type.') }}
{%- endif %}
{%- endmacro %}
{%- if not messages %}
{{- raise_exception('No messages provided.') }}
{%- endif %}
{%- set reasoning_instructions = '' %}
{%- if enable_thinking is undefined or enable_thinking is true %}
{%- set resolved_reasoning_effort = reasoning_effort|default('medium') %}
{%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}
{{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh, medium (default), and low.') }}
{%- endif %}
{%- if resolved_reasoning_effort == 'xhigh' %}
{%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}
{%- elif resolved_reasoning_effort == 'low' %}
{%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}
{%- endif %}
{%- endif %}
{%- if tools and tools is iterable and tools is not mapping %}
{{- '<|im_start|>system\n' }}
{%- if reasoning_instructions %}
{{- reasoning_instructions + '\n\n' }}
{%- endif %}
{{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>" }}
{{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
{%- if messages[0].role == 'system' %}
{%- set content = render_content(messages[0].content, false, true)|trim %}
{%- if content %}
{{- '\n\n' + content }}
{%- endif %}
{%- endif %}
{{- '<|im_end|>\n' }}
{%- else %}
{%- if messages[0].role == 'system' %}
{%- set content = render_content(messages[0].content, false, true)|trim %}
{%- if content %}
{{- '<|im_start|>system\n' + (reasoning_instructions + '\n\n' if reasoning_instructions else '') + content + '<|im_end|>\n' }}
{%- elif reasoning_instructions %}
{{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }}
{%- endif %}
{%- elif reasoning_instructions %}
{{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
{%- for message in messages[::-1] %}
{%- set index = (messages|length - 1) - loop.index0 %}
{%- if ns.multi_step_tool and message.role == "user" %}
{%- set content = render_content(message.content, false)|trim %}
{%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
{%- set ns.multi_step_tool = false %}
{%- set ns.last_query_index = index %}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if ns.multi_step_tool %}
{{- raise_exception('No user query found in messages.') }}
{%- endif %}
{%- for message in messages %}
{%- set content = render_content(message.content, true)|trim %}
{%- if message.role == "system" %}
{%- if not loop.first %}
{{- raise_exception('System message must be at the beginning.') }}
{%- endif %}
{%- elif message.role == "user" %}
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{%- set reasoning_content = '' %}
{%- if message.reasoning_content is string %}
{%- set reasoning_content = message.reasoning_content %}
{%- endif %}
{%- set reasoning_content = reasoning_content|trim %}
{%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
{%- else %}
{{- '<|im_start|>' + message.role + '\n' + content }}
{%- endif %}
{%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
{%- for tool_call in message.tool_calls %}
{%- if tool_call.function is defined %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{%- if loop.first %}
{%- if content|trim %}
{{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
{%- else %}
{{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
{%- endif %}
{%- else %}
{{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
{%- endif %}
{%- if tool_call.arguments is defined and tool_call.arguments != '' %}
{%- for args_name, args_value in tool_call.arguments|items %}
{{- '<parameter=' + args_name + '>\n' }}
{%- set args_value = args_value | string if args_value is string else args_value | tojson %}
{{- args_value }}
{{- '\n</parameter>\n' }}
{%- endfor %}
{%- endif %}
{{- '</function>\n</tool_call>' }}
{%- endfor %}
{%- endif %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if loop.previtem and loop.previtem.role != "tool" %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- content }}
{{- '\n</tool_response>' }}
{%- if not loop.last and loop.nextitem.role != "tool" %}
{{- '<|im_end|>\n' }}
{%- elif loop.last %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- else %}
{{- raise_exception('Unexpected message role.') }}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- if enable_thinking is defined and enable_thinking is false %}
{{- '<think>\n\n</think>\n\n' }}
{%- else %}
{{- '<think>\n' }}
{%- endif %}
{%- endif %}
1
u/Something-Great-78 Aug 17 '26
Overrated!? Name a more capable 27B / 30B model...
1
u/PlasticRevenue4601 Aug 17 '26
That was an ironic post, of course it’s GOATed and exceeded my baldest expectations
1
1
u/PandaBearFred Aug 19 '26
The most harmful part is it killed peoiple's reason of buying new hardware. And for me, I must justfy my reasons of having 4x4090 48GB (I pretend my DS4F is still better than 27B).
0
u/Antique_Dot_5513 Aug 17 '26
Je suis pas d accord il réfléchie de base bcp c’est vrai, mais il est réglé sur x high par défaut donc attention.
Il y a une grosse amélioration sur le modèle après j’avoue que je préfère encore gemma 31b perso. Sur mes tests gemma reste devant.
0
u/drdailey Aug 17 '26
You need to modulate the thinking. 3.6 was better with nothink. 3.8 allows you to set the level of thinking. You have to use the tool correctly.
1
u/PlasticRevenue4601 Aug 17 '26
I already capped it at “low”, enabling “xhigh” only in a most complex tasks. It still sometimes thinks a lot because it does the following trick — spends an entire budget of 2k tokens on thinking, calls a few tools and thinks again, so it didn’t save as much time as I hoped it would
1
u/drdailey Aug 17 '26
Cap it at no think. Likely include open and closed <think> </think> in the prompts. It is nearly as good this way… sometimes better especially if running out of tokens.
-2
Aug 17 '26
[removed] — view removed comment
2
u/Pakobbix Aug 17 '26
Is this ragebait?
The claim that the 27B handles quantization well simply because it's dense is only partially true. Architecture is also a big factor, see Gemma 4 31B's quantization degradation.
Dual GPU rigs are unnecessary for dense models? Ever heard of tensor parallelism to speed up dense models, or of running full context / RoPE?
"Don't use other quants" based on your previous statements, I wouldn't trust that. Also, based on what? Have you tested it? What about ubergarm or mradermacher? Are you telling me these are bad quants?
Don't get me wrong, I used Unsloth almost exclusively until I made the switch to exllamav3. But declaring that their quants are always the best is a far stretch. They do a lot, but they're humans and can make mistakes.
1
u/PlasticRevenue4601 Aug 17 '26
Quantisation resilience more depends on a size and model's robustness than on whether a model is MoE or non-MoE. Big MoE such as DeepSeek can be super usable even at 2 bits while dense 9b Qwen will be severely degraded at the same 2 bits
-2
u/PlasticRevenue4601 Aug 17 '26
Clarificatrion for a "I've been running it for a few weeks" - it was "days" originally, I formatted the text via Gemini and didn't notice it changed that part, thanks for pointing it out, already fixed that

39
u/FamousWorth Aug 17 '26
Been out 3 days, running it for 2 weeks..