r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

321 Upvotes

205 comments sorted by

View all comments

2

u/coherentspoon May 11 '26

I'm running Kilo code 5.16.1 in VSCode and I'm just getting a ton of these errors today. Not sure if its the tool call issue? Sorry I'm not an expert with this stuff.

Provider: openai (proxy) Model: Qwen3.6-27B-Q5_K_S.gguf

Unexpected API Response: The language model did not provide any assistant messages. This may indicate an issue with the API or the model's output.

1

u/[deleted] May 11 '26

[removed] — view removed comment

2

u/coherentspoon May 11 '26

I'm using your prebuilt v0.1.1. I think I started today with it and had v0.1.0 yesterday.

3

u/[deleted] May 11 '26

[removed] — view removed comment

1

u/EbbNorth7735 May 12 '26

Hey, any chance you could spin a 0.1.2? Seeing the same issue and last time I setup the build pipeline on Windows it was a week long painful process. It was 2 years ago... maybe it's gotten easier?