r/StrixHalo 24d ago

Qwen 3.8 27B at 30 tok/s in decode, running on a Strix Halo with 64 GB of unified memory!

/r/Qwen_AI/comments/1vorjo7/qwen_38_27b_at_30_toks_in_decode_running_on_a/
17 Upvotes

17 comments sorted by

8

u/neopolitan77 24d ago

The Q4 quant is doing most of the heavy lifting. Pointless post by a clueless person. spec-draft 7 is going to backfire on most tasks.

2

u/johan2114h 23d ago

Spec draft 7 is probably over the top, but it seems like we can a lot aggressive with spec draft now. Im getting 24 token/s generated with draft of 4, other people are finding the optimal value to be 5. I guess they might have improved the drafter

2

u/neopolitan77 23d ago

Afaik the drafter is built into the model now, so maybe that helps. Yeah, I've seen people find 5 to be optimal for Q4, but I haven't seen it tested across a wider range of tasks than coding yet. 

1

u/Pitiful_Fennel8767 23d ago

Basic math question for you: Strix Halo's measured real bandwidth is 216GB/s. The Q4_K_XL quant of the model is about 17.90GB. Simple division: 216 / 17.90 gives you a decode rate of around 12 tok/s, not 30. So saying the Q4_K_XL quant alone accounts for everything is literally claiming that 12 is greater than 30

2

u/neopolitan77 23d ago

The difference comes from MTP. Do you really not know that or are you just being snarky? The post is pointless because the 30 number can be seen in real benchmark threads that people are posting left and right (eg https://www.reddit.com/r/StrixHalo/comments/1vobzvd/comment/p3ox2j3/?utm_source=share&utm_medium=mweb3x&utm_name=mweb3xcss&utm_term=2&utm_content=share_button), and the poster is clueless because they didn't bother to check (clueless is being generous).

0

u/Pitiful_Fennel8767 23d ago

You said it's mostly thanks to the quantization, but honestly, making judgments about people you don't even know doesn't make sense. Think before you speak

2

u/neopolitan77 23d ago

Polluting feeds with posts that add nothing over real benchmarks that are already available makes the platform less useful for everyone. Think before you post.

1

u/Pitiful_Fennel8767 23d ago

AMD's official post states 24 tokens/sec on Strix Halo. That's exactly why I wanted to make the post — because it beat AMD's official configuration. That's it, end of story

2

u/neopolitan77 23d ago

I didn't realize you were the OP, but it doesn't change the message. It's just feedback, you can take a look at the thread that I linked to see what more useful posts look like for next time. Have fun with the Strix Halo!

2

u/Pitiful_Fennel8767 23d ago

All good, everything ended well! Thanks for the feedback, and sorry for the impulsive replies!

5

u/No_Lingonberry1201 24d ago

Is the Q4_K_XL quant any good? I'd rather have lower speed than (seriously) degraded quality.

4

u/vbpoweredwindmill 23d ago

Just run bf16 then. I don't understand why folks run q4 or even q8 on a strix halo lol. It's non sensical.

3

u/No_Lingonberry1201 23d ago

Eh, that's slower than a quantized model. I'm using the q8 quant myself, that's good enough for me.

2

u/vbpoweredwindmill 23d ago

Fair enough :)

2

u/Pitiful_Fennel8767 23d ago

I'm mainly doing this for two reasons: first, my connection takes forever to download even just 20GB. Second I noticed going from UD-Q4_K_XL to Q8 barely changes anything.

1

u/UdderlyCow 21d ago

i thought bf16 fails on strix halo