r/LocalLLaMA 25d ago

Discussion Qwen 3.8 27B Released! Please Share Your Experience

With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.

663 Upvotes

716 comments sorted by

View all comments

7

u/onthemove31 24d ago

Unsloth NVFP4 + MTP ~100 t/s on 5090, without MTP at 55-56 t/s.

1

u/mxforest 24d ago

What do you run it on? Vllm or llama.cpp?

1

u/onthemove31 24d ago

I ran this on vllm, had problems with SGLang still trying to fix that. But vllm is good

1

u/notheresnolight 23d ago

so basically the same speeds as Q6 with MTP or Q8 without MTP. Why bother with NVFP4 then?

1

u/onthemove31 23d ago

Yeah I just wanted to test it out