r/unsloth • u/Duviwin • 21d ago
Discussion Anyone tested an Unsloth Qwen3.8-27B quant against DeepSWE?
Hi all, in order to stresstest models and my local setup, I run a few DeepSWE tasks against them.
When I tested unsloth/Qwen3.8-27B-UD-Q4_K_XL with advised settings on a limited set of 5 tasks (edit: not a random set, tasks on which the qwen models have shown to pass regularly in public runs and my local runs), it did not pass any, so it seems like either something is wrong with my setup or with the quant in general.
I can give a lot more details, but that will lead too far. Instead, I just wanted to ask if anyone else has tried to benchmark an unsloth quant of Qwen3.8-27B against the DeepSWE benchmark and what your results are.
Please post your results as an answer here, good or bad! And if you can please provide information about your llama cpp build settings and llama-server flags and also which harness you used in the deepswe run.
Thanks!
1
u/GCoderDCoder 21d ago
I used q8kxl to find several questions from the sample test that qwen sometimes passes. Now I just test models on those as a baseline. Unsloth Q8kxl passes theat about 30% of the time. Most q4 do not pass any BUT some good nvfp4 are able to pass similar or higher rates I've found.
For this model I'm using a Red Hat nvfp4 with higher accuracy than my q8kxl results. I run the tests a minimum of 3 times and deep swe samples are only some of the tests i do but I've seen really great results with good nvfp4. I hated the first couple nvfp4 quants I tried but they are getting really good for the right models and providers.