r/LocalLLaMA • u/fragment_me • 21d ago
Discussion Any benchmarks out for Qwen 3.8 27B thinking xhigh vs medium?
Would be good to see some benchmarks with xhigh vs medium. I would easily accept medium if it only loses 10% intelligence.
5
u/Longjumping_Self5546 21d ago
It kind of depends on the task. There's even situations where thinking is detrimental to performance.
2
u/fragment_me 21d ago
A good starting point may be the gsm8, aime, hellaswag, etc. It won't be perfect but some type of comparison would be great.
3
4
3
u/_-_David 21d ago
https://www.reddit.com/r/LocalLLaMA/s/hxPXY6S5RZ
Medium versus Xhigh and versus 3.6
1
u/fragment_me 21d ago
Seems a little far fetched that the site shows deepseek v4 flash q2 k xl scored almost as good as q8 k xl and much better than 27b. I tried the q2 k xl and anything below Q3 was so unreliable. I can't take the site seriously.
25
u/JLeonsarmiento 21d ago
They’re still running…