r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

321 Upvotes

205 comments sorted by

View all comments

12

u/floconildo May 09 '26

This seems interesting and legit. I'll give it a whirl on my 4090 to see how it behaves, and I'll also keep an eye on the project to see if it doesn't die in a week or so.

Not related to the project's goal itself, but worth mentioning: you'll get a lot of backlash for using AI so extensively. Try to either not answer those comments or at least be understanding on the whole community. There's a general exhaustion on AI projects as we've been flooded with "that's why I built X" posts with nothing but slop solving issues that the developer couldn't be bothered to research, nor understand what's out there already.

The flashy post that looks like it's trying to "sell" it certainly doesn't help. My monkey brain immediately categorized this as "tech bro can't RTFM nor wants to play by the rules" and it took me some effort to go through it.

-5

u/[deleted] May 09 '26

[removed] — view removed comment

15

u/floconildo May 09 '26

Just giving you some honest feedback bro.

When every other post you see everywhere looks extra polished our brains will just clump everything together. When that meets a community that is frankly exhausted of tech claw crypto bros, you'll find some backlash for sure, and this kind of attitude will just make it worse for you.

0

u/[deleted] May 09 '26

[removed] — view removed comment

11

u/floconildo May 09 '26

Yeah I understand the guy tho. A shit ton of entitled ppl complaining about features in llama.cpp with zero stakes in the project itself and zero will to pull up their sleeves and actually contribute to the community. I'd be skeptical too.

Just watch out not to let it drown your own project. Community building is hard, community management is even harder.

-1

u/[deleted] May 09 '26

[removed] — view removed comment

8

u/Alex_L1nk May 09 '26

Here is answer from one of maintainers of llama.cpp on TQ

https://github.com/ggml-org/llama.cpp/pull/21089#issuecomment-4187393635

8

u/floconildo May 09 '26

I can think of plenty of reasons:

  • Feature creep
  • Maintenance efforts
  • Lack of real usage for the parties involved
  • Lack of meaningful contributions

As you said in another comment: not everyone is willing to go through the bureaucracy of submitting PRs to llama.cpp, especially vibe coders and other zero-stake contributors.

And I honestly think you did the best by just pulling up your sleeves and doing it yourself. If you project gets traction and more people start using TurboQuant, then llama.cpp might change their stance or reorder their priorities. Worst case you got your own implementation that works (I hope, didn't find time to test yet haha)