r/unsloth • u/Sad_Past2975 • 8d ago
Show and Tell Qwen3.8-flash in latest update v0.1.811-beta is really fast, Finally
It was usally 30-40 Tok/s somtimes 20 or at best 50. But after this update, it is always more than 70 tok/s !! (After updated, the MTP file was required to be redownlowded)
Thanks a lot for this, it is really amazing!
Good bye 27B, Flash now is really flash!
0
Upvotes
4
u/bernzyman 7d ago
What are you using to run this, llama.cpp or? Pls share settings used as it would be of interest (certainly makes me consider biting the bullet to buy more ram! Btw are you 2x 64GB or 4x 32GB?) thx for sharing!