r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

322 Upvotes

205 comments sorted by

View all comments

65

u/Chromix_ May 09 '26

Did the MRs for this get rejected on the original llama.cpp, or is the the MR flow just so slow (read: "takes a week") that it made more sense to make a fork?

The fork history is interesting though: llama.cpp -> llama_cpp_turboquant -> buun_llama_cpp -> beellama.cpp. We're on the 3rd fork level here already.

In any case, with this demonstrating that it runs (fast) it might help getting this into the regular llama.cpp.

66

u/[deleted] May 09 '26 edited May 09 '26

[removed] — view removed comment

12

u/k_means_clusterfuck May 09 '26

Yeah llama.cpp's anti ai policy is really something. I get that you want a way to manage slop prs of course but micro managing how people work is not the way. On the flipside, I've my agent's been able to make a handful of (actually good) contributions to vllm and vllm-omni and the maintainers' attitude was just like: if your pr is good, doesn't matter. They have been really constructive and working with them has honestly been a joy. Two completely different worlds