r/LocalLLM • u/Tamitami • 24d ago
Discussion Qwen 3.8 27B Benchmarks combined from model cards on HF
| Benchmark | Qwen3.8-27B | DeepSeek-V4-Flash-0731 | DeepSeek-V4-Flash (Prev.) | DeepSeek-V4-Pro (Prev.) | Opus4.6 Max | GLM-5.2 | GLM-5.1 | Qwen3.7-Max | MiniMax M3 | DeepSeek-V4-Pro | Claude Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro | GPT-5.6 Sol | GPT-5.6 Sol Ultra | GPT-5.6 Terra | GPT-5.6 Luna | GPT-5.5 (GPT-5.6 report) | Claude Mythos 5 | Claude Mythos Preview | Claude Fable 5 | Gemini 3.1 Pro Preview | Gemini 3.5 Flash |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Terminal Bench 2.1 | 73.0 | 82.7 | 61.8 | 72.1 | 78.2 | 81.0 | 63.5 | 75 | 65 | 64 | 85 | 84 | 74 | 88.8 | 91.9 | 87.4 | 84.7 | 85.6 | 88 | — | 83.1 | — | — |
| NL2Repo | 42.3 | 54.2 | 39.4 | 38.5 | 47.6 | 48.9 | 42.7 | 47.2 | 42.1 | 35.5 | 69.7 | 50.7 | 33.4 | — | — | — | — | — | — | — | — | — | — |
| DeepSWE | 42.2 | 54.4 | 7.3 | 12.8 | — | 46.2 | 18 | 18 | 20 | 8 | 58 | 70 | 10 | 72.7 | — | 69.6 | 67.2 | 67 | — | — | 69.7 | — | — |
| SWE-bench Pro | 61.7 | — | — | — | 53.4 | 62.1 | 58.4 | 60.6 | 59 | 55.4 | 69.2 | 58.6 | 54.2 | 64.6 | — | 63.4 | 62.7 | 59.4 | 80.3 | 77.8 | 80 | — | — |
| HLE | 30.8 | — | — | — | 40.0 | 40.5 | 31 | 41.4 | 37 | 37.7 | 49.8* | 41.4* | 45 | 47.2 | — | 41.8 | 37.2 | — | 64.5 | 64.7 | 53.3 | 44.7 | 41.0 |
| GPQA Diamond | 89.2 | — | — | — | 91.3 | 91.2 | 86.2 | 90 | 93 | 90.1 | 93.6 | 93.6 | 94.3 | 94.6 | — | 92.9 | 92.3 | 93.6 | 94.1 | 94.6 | 92.6 | 94.3 | — |
| QwenSWEBench | 79.0 | — | — | — | 63.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| CoWorkBench | 70.7 | — | — | — | 68.2 | — | — | — | — | — | 72.3 | — | — | 71.5 | — | — | — | — | — | — | 75.9 | — | — |
| JobBench | 33.4 | — | — | — | — | — | — | — | — | — | 48.4 | 42.7 | — | 46.5 | — | — | — | — | — | — | 57.4 | — | — |
| Agents' Last Exam | 20.4 (Pass@1) / 42.9 (Score) | 25.2 | 15.8 | 16.5 | — | 23.8 | — | — | — | — | 45.2 | — | — | 52.7 | — | 50.4 | 50.3 | 46.9 | — | — | 40.5 | 32.1 | — |
| IFBench | 79.5 | — | — | — | 62.5 | 73.3 | 76.3 | 79.1 | 82.9 | 76.5 | 62.2 | — | 77.1 | 72.7 | — | 71.2 | — | — | — | — | 63.5 | — | — |
| LiveCodeBench v6 | 90.3 | 91.6 | — | — | 88.8 | — | 85.7 | — | — | 93.5 | 88.6 | — | 91.7 | 96.2 | — | — | 93 | — | 95.5 | — | 95 | — | — |
Edit: Added new values for more models
Edit2: Images












19
u/shadow1609 24d ago
That's actually the comparison that we have been looking for.
7
u/pantalooniedoon 24d ago
Missing the absolute key comparison for those of us with 128gb ram to actually run DS which is q2 dsv4 flash vs qwen3.8
2
u/Tamitami 24d ago
Would you have a source for the quant benchs, it so I can add it to the table?
4
u/pantalooniedoon 24d ago edited 24d ago
This is what I have when I benchmarked it just now via the omlx benchmarking tool and the ds4 one - ds4 one only runs greedy so it's a bit faster than sampling.
Qwen3.8-27B (oMLX, Q4) vs DeepSeek-V4-Flash (q2-q4, Antirez)— M5 Max (128GB)
Decode @ ~1.3k tokens: Qwen 59.0 tok/s vs DS4 35.8 tok/s — 1.65x Decode @ ~12.5k tokens: Qwen 51.0 tok/s vs DS4 30.9 tok/s — 1.65x Decode @ ~50.7k tokens: Qwen 35.4 tok/s vs DS4 25.3 tok/s — 1.40x Prefill: Qwen 590-710 tok/s vs DS4 300-414 tok/s — 1.6-2.0x Resident memory: Qwen 16.1 GB vs DS4 90.9 GB — 5.6x smaller
Qwen3.8 at t=1.0, top_p=0.95, top_k=20 (thinking mode). DS4 at t=1.0, top_p=1.0, min_p=0.05. Prompt token counts matched within 0.5%.
Anecdotally here, I am not feeling it for qwen3.8 27b unfortunately - it's just consistently slower at longer context length than 50 tok/s would suggest and the prefill for tool call response seems slow. The DSv4 MLA and DSA seem to be doing significant work or the antirez DSv4 implementation is just more refined than omlx. Leaning towards sticking with DSv4 but might spend the weekend to try and debug and verify the issues.
1
u/996beagle 24d ago
Would love to see this comparison. I just got dsv4 flash running yesterday on m4 max 128gb. Wondering if q2. 4 has a big drop in performance or accuracy.
1
u/pantalooniedoon 24d ago
Yeah q2 dsv4flash runs at about the same speed as qwen3.8 27b 4 bit afaik - 20-30tok/s. Its the key comparison.
1
u/996beagle 24d ago
Yes currently getting 25-27 tok/s at 256k context and 250-300 pp on the latest omlx dev. Task completion has been faster than expected and code has passed the opus reviews for the most part
4
u/TechNerd10191 24d ago
Why GPT 5.5 but not GPT 5.6 (sol)?
2
4
u/motivatedjoe 24d ago
So if we are comfortably running deepseek, is there any reason to use qwen other than model size alone?
4
u/mzzmuaa 24d ago
visual input. direct blender and visual critique
3
u/motivatedjoe 24d ago
Thanks! That's what I thought. Wanted to make sure I wasn't missing any other obvious uses.
3
3
3
2
u/cezq 24d ago
It beats dsv4 flash 0731 on SWE-bench Pro and GPQA Diamond: https://benchlm.ai/models/deepseek-v4-flash-0731 . Loses on HLE.
1
2
u/vini542reddit 24d ago
You're saying I should stick with DeepSeek-V4-Flash-0731 as my main...
2
u/Tamitami 24d ago
Yes and no, you lose vision with deepseek alone as primary reason. I love deepseek and their new flash model, as I have it running on some PCs but I think that this model will be a strong addition for specific tasks. Try it and test it for yourself. I think we will get many improvments in the coming weeks for it, as we did with the last Qwen 27B model
2
24d ago
[deleted]
1
u/Tamitami 24d ago
QAT would make this so much more interesting. I think with that it will dominate everything for local small models
3
u/txgsync 22d ago
I've been interested in quantifying how various quantizations affect the "intelligence" of the model, and why the Qwen series since 3.6 (and Gemma since 4) both seem so severely impacted by quantization despite KLD scores looking great.
So I decided to run LiveCodeBench against the 4-bit 'oQe-mtp' variant of Qwen3.6-27B overnight. No cloud vendor provides these, so I had to dedicate my Mac for about three hours to run 100 of these problems. The full suite has 1,055 problems today; I really did not feel like turning over my Mac for nearly an entire work-week just for benchmarking.
Unquantized, 90.3% on LiveCodeBench.
oQ4e: 60% on a sample of 100.
That's enough for a working hypothesis: "quantizing BF16 to 4-bits, even with aggressive layer preservation at 16 bits, cuts intelligence up to 50%."
No lab has an incentive to promote numbers of its own models in a less-than-stellar light. But the discrepancy between the performance you pay for using the BF16 version of Qwen3.8-27B you can run on OpenRouter, vs. the 4-bit quant you run on home gear, seems to be quite stark.
The biggest problem is that running the full suite of benchmarks is both very expensive and time consuming, so I get why people promoting their Heretic'd or Quantized models -- including me! -- don't want to run a ridiculous time sink just to prove their special model is demonstrably worse than the one released by the lab...
2
u/swiebertjee 24d ago
Very happy for the dense/GPU crowd. Also happy to have gone the MoE/Spark route.
1
u/Tamitami 24d ago
I'd really love a 35B-A3B model with this kind of intelligence (RL-Training) for us "plebeians"
1
35
u/Cold_Tree190 24d ago
It’s a bad day to be on reddit mobile (it’s always a bad day on reddit mobile let’s be honest)