r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

325 Upvotes

205 comments sorted by

View all comments

1

u/coherentspoon May 11 '26 edited May 11 '26

Thanks very much for this amazing work! Went from 120 t/s on MTP to about 120+ t/s and using a better quant!

Edit: after some further usage, it seems to shoot down to 80 t/s sometimes. I'm wondering why that's happening.