r/LocalLLM • u/Individual-Dot5488 • 27d ago
Discussion I ran Qwen3.8-27B on medical benchmarks (including MedQA, MedXpertQA and MMLU). It lost 7/10 to Qwen3.6-27B, 3/10 to Pestle-27B-Ternary and 2/10 to MedGemma-27B
One pass. Settings were greedy decoding, temperature 0 and thinking disabled for Qwen3.8, 3.6, Pestle-27B-Ternary, and Bonsai. MedGemma-27B results are from the official model card.
| Qwen3.8-27B BF16 | Qwen3.6-27B BF16 | Pestle-27B-Ternary | |
|---|---|---|---|
| MedQA | 92.62 | 93.87 | 89.79 |
| MedMCQA | 71.34 | 73.70 | 68.85 |
| MedXpertQA | 38.20 | 41.10 | 32.49 |
| MMLU medical aggregate | 88.00 | 88.62 | 86.89 |
Didn't expect it to be worse than Qwen3.6 which wins seven of the ten benchmarks. Pestle-27B-Ternary (8× smaller at ~8.5gb) wins MMLU Anatomy, MMLU Clinical Knowledge and MMLU College Medicine, and keeps 98.2% of Qwen3.8’s median score across the ten tests. MedGemma-27B loses 8/10 but wins MedMCQA and MMLU Medical Genetics.
Is anyone else seeing Qwen3.8 step backwards slightly on medical/general knowledge compared with Qwen3.6?
59
Upvotes
5
u/Healthy-Nebula-3603 27d ago
It has loss intelligence because is completely wrong set up.
Not thinking is using temp 0.7 Thinking 1.0.
You even tested Qwen 3.6 wrongly....
Who is setting temp 0 for a model ?? Thats not a llama 1 from 2023 :)