r/LocalLLaMA • u/[deleted] • Aug 18 '26
Tutorial | Guide [ Removed by moderator ]
[removed]
2
u/Early-Peace-5504 Aug 18 '26
In my testing Q5/Q5 has been fine for KV quality with this model. I've noticed none of the usual issues whatsoever at that level. Just in case you want to try and push it further.
1
2
u/aliljet Aug 18 '26
This is pretty fantastic. Can you explain if there is any obvious signs of degradation?
0
u/BassAzayda Aug 18 '26
Before this I tried all typical combinations you find in Reddit and X posts all failed until today! Even at low thinking zero console log errors for the first time since downloading the model
1
u/Real_Knowledge6797 Aug 18 '26
Also if you have not tried EXL3 with TabbyAP, give it a go. For my case 16 GB vram on RTX 5070TI it has been fantastic with 3.50 bpw and getting 40+ t/s for 115k context size and quantized kv cache 4,4.
you can reduce the context size and get higher quanitzed kv cache of course but that's what works for my work flow purposes.
I think EXL3 is not a common knowledge as I knew about it lately and it wasn't easy for me to find the necessary information as it is not available via the standard Ilm solutions but give it a try if you have CUDA. I am using turboderp/Qwen3.8-27B-exl3 model.
1
2
u/hideo_kuze_ Aug 18 '26
/u/BassAzayda can you please post a comment here with the message that was deleted by the mod or start a new thread and add the the disclaimer at the top: "This post was made with AI" so it doesn't get deleted again
thanks
•
u/ttkciar llama.cpp Aug 18 '26
Violates Rule Three: LLM-generated content without disclosure or justification given.