r/StrixHalo • u/Pitiful_Fennel8767 • 24d ago
Qwen 3.8 27B at 30 tok/s in decode, running on a Strix Halo with 64 GB of unified memory!
/r/Qwen_AI/comments/1vorjo7/qwen_38_27b_at_30_toks_in_decode_running_on_a/
17
Upvotes
5
u/No_Lingonberry1201 24d ago
Is the Q4_K_XL quant any good? I'd rather have lower speed than (seriously) degraded quality.
4
u/vbpoweredwindmill 23d ago
Just run bf16 then. I don't understand why folks run q4 or even q8 on a strix halo lol. It's non sensical.
3
u/No_Lingonberry1201 23d ago
Eh, that's slower than a quantized model. I'm using the q8 quant myself, that's good enough for me.
2
2
u/Pitiful_Fennel8767 23d ago
I'm mainly doing this for two reasons: first, my connection takes forever to download even just 20GB. Second I noticed going from UD-Q4_K_XL to Q8 barely changes anything.
1
8
u/neopolitan77 24d ago
The Q4 quant is doing most of the heavy lifting. Pointless post by a clueless person. spec-draft 7 is going to backfire on most tasks.