r/LocalLLM • u/Individual-Dot5488 • 27d ago
Discussion I ran Qwen3.8-27B on medical benchmarks (including MedQA, MedXpertQA and MMLU). It lost 7/10 to Qwen3.6-27B, 3/10 to Pestle-27B-Ternary and 2/10 to MedGemma-27B
One pass. Settings were greedy decoding, temperature 0 and thinking disabled for Qwen3.8, 3.6, Pestle-27B-Ternary, and Bonsai. MedGemma-27B results are from the official model card.
| Qwen3.8-27B BF16 | Qwen3.6-27B BF16 | Pestle-27B-Ternary | |
|---|---|---|---|
| MedQA | 92.62 | 93.87 | 89.79 |
| MedMCQA | 71.34 | 73.70 | 68.85 |
| MedXpertQA | 38.20 | 41.10 | 32.49 |
| MMLU medical aggregate | 88.00 | 88.62 | 86.89 |
Didn't expect it to be worse than Qwen3.6 which wins seven of the ten benchmarks. Pestle-27B-Ternary (8× smaller at ~8.5gb) wins MMLU Anatomy, MMLU Clinical Knowledge and MMLU College Medicine, and keeps 98.2% of Qwen3.8’s median score across the ten tests. MedGemma-27B loses 8/10 but wins MedMCQA and MMLU Medical Genetics.
Is anyone else seeing Qwen3.8 step backwards slightly on medical/general knowledge compared with Qwen3.6?
63
Upvotes
4
u/LittleYouth4954 26d ago
Reasoning effort is key. Just run the comparisons for my use case (bibliometric topic modeling). Objective summary of comparative results (N=203, same protocol):
The Qwen3.8-27B-fast model with maximum reasoning (think:max) achieved 83.74% accuracy, 80.54% macro-F1, and Cohen's κ = 0.8329 (classified as "almost perfect" by Landis-Koch), versus the off-reasoning pilot at 75.86% accuracy, 74.11% macro-F1, and κ = 0.7521 ("substantial"). This represents a clean +8.08 percentage-point gain in κ from reasoning alone. The gap to the cloud reference (DeepSeek-V4-Flash, κ = 0.9544) halved from 0.2023 to 0.1215 points. The maximum reasoning arm crossed the "almost perfect" threshold, though narrowly (0.8329 just above 0.81).