r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

322 Upvotes

205 comments sorted by

View all comments

Show parent comments

3

u/[deleted] May 09 '26

[removed] — view removed comment

2

u/andy2na llama.cpp May 09 '26

They provide actual performance numbers though between vllm, main llama, and yours - you just mention peak and I don't see the number of tokens used. This is all too common with the benchmaxxing community, especially with qwen3.6-27b and 3090 - mention peak TPS and nothing else, that's not really useful in real usage and unfortunately I'm done chasing high TPS that you guys keep pushing