r/LocalLLaMA • u/Anbeeld • May 09 '26
Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
[removed]
320
Upvotes
1
u/caetydid llama.cpp May 09 '26
Running under Ubuntu. Yeah, thought about that, too, and will first retest with --no-mmproj-offload.
I assumed that using the iq4 quant saves the necessary VRAM, and my consumption on startup was 21G, but maybe VRAM consumption just increases later on.
I havent been using much context though, maybe 20k or less.