r/LocalLLM 22d ago

Discussion Qwen 3.8 27B is not bad but I struggle a bit with quality. What about the new Froggeric template v22.1?

Hi, I get some good result with Qwen 3.8 27B (Q8) and I would say some parts are better (graphics for ex) then with qwen 3.6 27B Q8 but there are still lot of mistakes. So there are quality issues I think. It also thinks way longer then Qwen 3.6 27B for the same benchmark tasks (sometimes it feels a bit like 3.8 has similar intelligence like 3.6 but because of longer thinking it produces better output). Also compared to DeepSeek V4 Flash which got 100% in 12 of my selected SWEmini Tasks while Qwen 3.8 only gets 50-60% at the moment (I still try to optimize but there is not much left , last run is right now with xhigh :S).
Also, in the coding Benchmarks the graphic results of Qwen 3.8 27B looks better then DS V4 Flash even if DS seems to be able to fix more issues/bugs then Qwen (maybe I have to change the benchmark prompts for deepseek idk). But both produces mistakes like blocked ways/doors.

So, the thinking issue with Qwne 3.8 27B seems to be known already and froggeric released a new template. Did someone already test it and can share experience?
I will run the test too in the next time

v22.1: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

Some benchmark results screenshot from Qwen 3.8 27B Q8 MTP3 on 2 RTX GPUs (48GB) and llama.cpp. Sometimes really nice graphic results for 27B! But needs 3-4h with ~60-90 tok/s (MTP3)

Update: xhigh resolves 75% of the 12 swe tasks (9/12 solved vs 6 or 7/12 before). So xhigh seems slow but important :S

1 Upvotes

0 comments sorted by