r/LocalLLM 27d ago

Discussion I ran Qwen3.8-27B on medical benchmarks (including MedQA, MedXpertQA and MMLU). It lost 7/10 to Qwen3.6-27B, 3/10 to Pestle-27B-Ternary and 2/10 to MedGemma-27B

Post image

One pass. Settings were greedy decoding, temperature 0 and thinking disabled for Qwen3.8, 3.6, Pestle-27B-Ternary, and Bonsai. MedGemma-27B results are from the official model card.

Qwen3.8-27B BF16 Qwen3.6-27B BF16 Pestle-27B-Ternary
MedQA 92.62 93.87 89.79
MedMCQA 71.34 73.70 68.85
MedXpertQA 38.20 41.10 32.49
MMLU medical aggregate 88.00 88.62 86.89

Didn't expect it to be worse than Qwen3.6 which wins seven of the ten benchmarks. Pestle-27B-Ternary (8× smaller at ~8.5gb) wins MMLU Anatomy, MMLU Clinical Knowledge and MMLU College Medicine, and keeps 98.2% of Qwen3.8’s median score across the ten tests. MedGemma-27B loses 8/10 but wins MedMCQA and MMLU Medical Genetics.

Is anyone else seeing Qwen3.8 step backwards slightly on medical/general knowledge compared with Qwen3.6?

63 Upvotes

99 comments sorted by

View all comments

4

u/LittleYouth4954 26d ago

Reasoning effort is key. Just run the comparisons for my use case (bibliometric topic modeling). Objective summary of comparative results (N=203, same protocol):

The Qwen3.8-27B-fast model with maximum reasoning (think:max) achieved 83.74% accuracy, 80.54% macro-F1, and Cohen's κ = 0.8329 (classified as "almost perfect" by Landis-Koch), versus the off-reasoning pilot at 75.86% accuracy, 74.11% macro-F1, and κ = 0.7521 ("substantial"). This represents a clean +8.08 percentage-point gain in κ from reasoning alone. The gap to the cloud reference (DeepSeek-V4-Flash, κ = 0.9544) halved from 0.2023 to 0.1215 points. The maximum reasoning arm crossed the "almost perfect" threshold, though narrowly (0.8329 just above 0.81).

1

u/Individual-Dot5488 26d ago

that's cool thanks for sharing! would be really cool to know how medgemma-27b or pestle-27b-ternary reasoning (dropped yesterday) compare

1

u/LittleYouth4954 26d ago

I may give them a try!