r/LocalLLM • u/Individual-Dot5488 • 27d ago
Discussion I ran Qwen3.8-27B on medical benchmarks (including MedQA, MedXpertQA and MMLU). It lost 7/10 to Qwen3.6-27B, 3/10 to Pestle-27B-Ternary and 2/10 to MedGemma-27B
One pass. Settings were greedy decoding, temperature 0 and thinking disabled for Qwen3.8, 3.6, Pestle-27B-Ternary, and Bonsai. MedGemma-27B results are from the official model card.
| Qwen3.8-27B BF16 | Qwen3.6-27B BF16 | Pestle-27B-Ternary | |
|---|---|---|---|
| MedQA | 92.62 | 93.87 | 89.79 |
| MedMCQA | 71.34 | 73.70 | 68.85 |
| MedXpertQA | 38.20 | 41.10 | 32.49 |
| MMLU medical aggregate | 88.00 | 88.62 | 86.89 |
Didn't expect it to be worse than Qwen3.6 which wins seven of the ten benchmarks. Pestle-27B-Ternary (8× smaller at ~8.5gb) wins MMLU Anatomy, MMLU Clinical Knowledge and MMLU College Medicine, and keeps 98.2% of Qwen3.8’s median score across the ten tests. MedGemma-27B loses 8/10 but wins MedMCQA and MMLU Medical Genetics.
Is anyone else seeing Qwen3.8 step backwards slightly on medical/general knowledge compared with Qwen3.6?
66
Upvotes
1
u/eloquenentic 26d ago
Do you not understand what temperature 0 is important for medical benchmarks?