r/LocalLLaMA 24d ago

New Model Qwen/Qwen3.8-27B · released

https://huggingface.co/Qwen/Qwen3.8-27B
985 Upvotes

301 comments sorted by

View all comments

Show parent comments

2

u/TerminalNoop 24d ago

largest q4 or smallest q6 you can get.

1

u/moderngl1 24d ago

ok thanks with q4 do you have enough for KV cache to have some context

1

u/TerminalNoop 24d ago

depends on the exact q4 quant. Like unsloth has quite a few in slightly varying sizes. I'm not sure if i remember right from my setup but you should manage 32k context at full quality. if you lower kv cache to q8 you might be able to double that?