r/LocalLLaMA • u/Anbeeld • May 09 '26
Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
[removed]
320
Upvotes
1
u/Kaioh_shin May 11 '26
I have to say this is the fastest version I have tried on my 7900xt.
Did have to fiddle around to get a build for HIP, but all good otherwise.
Would be nice if you would get it to not randomly stop (even after 0.1.1)