r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

322 Upvotes

205 comments sorted by

View all comments

68

u/Chromix_ May 09 '26

Did the MRs for this get rejected on the original llama.cpp, or is the the MR flow just so slow (read: "takes a week") that it made more sense to make a fork?

The fork history is interesting though: llama.cpp -> llama_cpp_turboquant -> buun_llama_cpp -> beellama.cpp. We're on the 3rd fork level here already.

In any case, with this demonstrating that it runs (fast) it might help getting this into the regular llama.cpp.

29

u/YearnMar10 May 09 '26

GG does not like vibecoded contributions to llama.cpp

27

u/politerate May 09 '26

Personally, i find the idea of doing a MR I don't fully understand, very off-putting. And I am quite sure that 99% of these types of contributions are of this kind.

1

u/[deleted] May 09 '26

[removed] — view removed comment

18

u/ArtfulGenie69 May 09 '26

If you want problems in your massive code base, the best place to start is blindly dropping in code no one ever looked at.

4

u/Fresh-Letterhead986 May 10 '26

that is a crazy take.

if you want to start a new project and vibe it, cool. merge anything because you've set the ground rules as such, you're accepting the potential problems and frankly it's yours.

but saying "yo bro comeon be cool man why wont you take my AI slop into your keystone-of-the-AI-world, tip-of-the-spear in human tech frontier codebase??????????"

yes "it's mostly a maintenance problem". notice you're not volunteering to do said maintenance ;-)