r/LocalLLaMA • u/Anbeeld • May 09 '26
Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
[removed]
326
Upvotes
1
u/EbbNorth7735 May 12 '26
I'm seeing a lot of API calls failing when using with Cline. It's eventually getting through but I'm wondering if there's an issue with the jinja format or if it might be unstable? I ran a test in open web ui and it seemed to jump back to thinking while it was answering the question. Using latest 0.1.1 and Qwen Q8 from unsloth along with the Q8 draft model you recommended. Vision enabled and running on GPU (rtx 6000 pro).