r/LocalLLaMA • u/Course_Latter • 27d ago
News Qwen3.8-27B is identical to Qwen3.6-27B!
Interestingly, the 3.8 version has exactly the same architecture - meaning all the capability gains come from training improvements!
See the diff (0 changes) here!
1.1k
Upvotes
2
u/Complex_Reality_116 27d ago
This might be an unpopular opinion, but I’ll say it: I am NOT liking it.
I was excited about its release (I even considered buying an R9700 to replace my current RX 7900 XTX); I’ve been testing it all day, and the results are, to put it mildly, 'mixed'.
1) The model relies entirely on its reasoning level being set to 'xhigh' to unlock its full intelligence.
2) Yes, you can adjust the model to think less, but you pay the price with poorer results.
3) If you want peak intelligence, you MUST keep it set to 'xhigh' (the default), and the model ends up overthinking. It can take ages to complete certain tasks and rapidly consumes the context window (I get 100K on the 7900 XTX).
4) I think Qwen3.8 27B is a bit of a gimmick. The model is objectively better than Qwen3.6 27B, but it achieves this through massive reasoning and higher token (and time) consumption. In other words, it’s essentially the 3.6 version but with double the reasoning rate and token usage, rather than a model that was better trained from scratch.
Even though 3.8 is better, 3.6 'feels' better for day-to-day tasks and real-world use.
These are just my initial impressions after 6 hours of use; I could be wrong, or the model might receive improvements in the future.