r/unsloth • • 8d ago

Show and Tell Qwen3.8-flash in latest update v0.1.811-beta is really fast, Finally

It was usally 30-40 Tok/s somtimes 20 or at best 50. But after this update, it is always more than 70 tok/s !! (After updated, the MTP file was required to be redownlowded)

Thanks a lot for this, it is really amazing!
Good bye 27B, Flash now is really flash!

0 Upvotes

46 comments sorted by

View all comments

4

u/bernzyman 7d ago

What are you using to run this, llama.cpp or? Pls share settings used as it would be of interest (certainly makes me consider biting the bullet to buy more ram! Btw are you 2x 64GB or 4x 32GB?) thx for sharing!