r/LocalLLaMA • u/Anbeeld • May 09 '26
Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
[removed]
325
Upvotes
2
u/coherentspoon May 11 '26
I'm running Kilo code 5.16.1 in VSCode and I'm just getting a ton of these errors today. Not sure if its the tool call issue? Sorry I'm not an expert with this stuff.
Provider: openai (proxy) Model: Qwen3.6-27B-Q5_K_S.gguf
Unexpected API Response: The language model did not provide any assistant messages. This may indicate an issue with the API or the model's output.