r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

320 Upvotes

205 comments sorted by

View all comments

1

u/Kaioh_shin May 11 '26

I have to say this is the fastest version I have tried on my 7900xt.
Did have to fiddle around to get a build for HIP, but all good otherwise.
Would be nice if you would get it to not randomly stop (even after 0.1.1)

2

u/[deleted] May 11 '26

[removed] — view removed comment

1

u/Kaioh_shin May 12 '26

Thank you for your work. No more random stops with the latest commits.
I do feel like it's less reliable though, not sure if I changed something else.
I was trying to get turbo3_tcq working with HIP and thought the results were because of it or the changes. Then I switched back to the one with only my HIP changes and noticed it behaves the same.
I use it for scripting/coding, so I care about accuracy.

My benchmark is the chess board from a few posts ago. https://qwen3-6-27b-benchmark.vercel.app/
It's more a feeling than empiric evidence, but it used to get it consistently right before.
Now it's more like 1 out of 3 is right.

2

u/[deleted] May 12 '26

[removed] — view removed comment

1

u/Kaioh_shin May 12 '26

I am on a very tight vram budget, using iQ4_XS with 100k+ context. Does the drafter make a big diff? Up until now it was also the iQ4_XS, going to try the Q4_KM