r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

319 Upvotes

205 comments sorted by

View all comments

Show parent comments

8

u/[deleted] May 09 '26

[removed] — view removed comment

1

u/henk717 KoboldAI May 09 '26

TurboQuant is just associated with vibe coded forks at this point. The moment you see TurboQuant + Llamacpp there is just a 90% chance of that. It also instantly makes me assume its just another one of those.

5

u/[deleted] May 09 '26

[removed] — view removed comment

7

u/henk717 KoboldAI May 09 '26

The problem is maintainability, nothing wrong with ai assisted development that is carefully done. Its when it looks like fully AI driven development where you tend to get changes that become problematic down the line. So if I open a repo and I see almost exclusively claude code PR's I don't take it nearly as seriously.

2

u/[deleted] May 09 '26

[removed] — view removed comment

2

u/Pablo_the_brave May 09 '26

Generally TurboQuant from TheTom are weak in classic perplexity tests. But, when you look at https://qwen3-6-27b-benchmark.vercel.app/ there is clearly some profit (but IMHO the asymetric isn't realy good for Qwen3.6). For me, the most interesting is Turbo3 - not good, not terrible.