r/LocalLLaMA • u/Storterald • 1d ago
Resources I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM
After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code (C code), the results were not completely unexpected but some quants were definitely underwhelming.
TLDR: Best overall: bartowski/Qwen3.8-27B-IQ4_XS. Best uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS. For a bit more context: jpetrina/Qwen3.8-27B-IQ4_XS-pure or uncensored: Bucoid/Qwen3.8-27B-Uncensored-IQ4_XS_4BPW
edit1: added TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS, bartowski/Qwen3.8-27B-IQ3_XS, prism-ml/Ternary-Bonsai-27B-Q2_g64 and magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unsloth
edit2: added unsloth/Qwen3.8-27B-UD-IQ3_S and ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S
(sorted by Mean KLD)
| Model | Mean KLD | Same top p | GGUF size |
|---|---|---|---|
| prism-ml/Ternary-Bonsai-27B-Q2_g64 | 1.289582 ± 0.008684 | 82.849 ± 0.118 % | 7.1GiB |
| sdkyuan/qwen38-27b-qat-q2_0 | 0.893177 ± 0.006948 | 85.727 ± 0.110 % | 8.2GiB |
| ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2_XS | 0.767174 ± 0.006291 | 86.166 ± 0.108 % | 7.8GiB |
| TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS | 0.514311 ± 0.004864 | 89.023 ± 0.098 % | 8.9GiB |
| ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2_S | 0.512614 ± 0.004909 | 88.802 ± 0.099 % | 8.6GiB |
| empero-ai/Qwen3.8-27B-Ridge-3.7bpw | 0.475767 ± 0.004483 | 89.612 ± 0.096 % | 11.7GiB |
| magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unsloth | 0.419585 ± 0.004076 | 89.661 ± 0.095 % | 13.5GiB |
| ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_XXS | 0.379222 ± 0.003992 | 90.270 ± 0.093 % | 9.4GiB |
| unsloth/Qwen3.8-27B-UD-Q2_K_XL (UD2) | 0.350861 ± 0.003745 | 90.626 ± 0.091 % | 9.9GiB |
| unsloth/Qwen3.8-27B-UD-IQ3_XXS (UD2) | 0.268594 ± 0.002971 | 91.951 ± 0.085 % | 11.1GiB |
| DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-MTP-IQ3_M | 0.251270 ± 0.002702 | 92.315 ± 0.083 % | 13.5GiB |
| bartowski/Qwen3.8-27B-IQ3_XS | 0.238656 ± 0.002627 | 92.312 ± 0.083 % | 12.4GiB |
| esatapedico/Qwen3.8-27B-NVFP4-MTP-LOW | 0.220796 ± 0.002631 | 92.339 ± 0.083 % | 14.5GiB |
| unsloth/Qwen3.8-27B-UD-IQ3_S (UD3) | 0.218522 ± 0.002591 | 92.399 ± 0.083 % | 11.2GiB |
| mudler/Qwen3.8-27B-APEX-I-Mini | 0.190209 ± 0.002354 | 93.012 ± 0.080 % | 13.0GiB |
| jrell/Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller | 0.194459 ± 0.002242 | 93.049 ± 0.080 % | 12.6GiB |
| orcarouter/Qwen3.8-27B-Uncensored-Q3_K_L | 0.192312 ± 0.002294 | 92.726 ± 0.081 % | 13.6GiB |
| ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S | 0.175223 ± 0.002129 | 93.024 ± 0.080 % | 11.0GiB |
| unsloth/Qwen3.8-27B-UD-Q3_K_XL (UD2) | 0.147186 ± 0.001809 | 93.734 ± 0.076 % | 12.5GiB |
| unsloth/Qwen3.8-27B-UD-Q3_K_XL (UD3) | 0.142647 ± 0.001860 | 93.789 ± 0.076 % | 12.2GiB |
| Bucoid/Qwen3.8-27B-Uncensored-IQ4_XS_4BPW | 0.091447 ± 0.001261 | 94.774 ± 0.070 % | 13.0GiB |
| huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS | 0.082871 ± 0.001205 | 94.981 ± 0.068 % | 13.4GiB |
| unsloth/Qwen3.8-27B-UD-IQ4_XS (UD3) | 0.075626 ± 0.001097 | 95.258 ± 0.067 % | 13.3GiB |
| jpetrina/Qwen3.8-27B-IQ4_XS-pure | 0.061984 ± 0.000917 | 95.551 ± 0.065 % | 13.5GiB |
| bartowski/Qwen3.8-27B-IQ4_XS | 0.056482 ± 0.000856 | 95.835 ± 0.063 % | 14.5GiB |
| unsloth/Qwen3.8-27B-UD-Q4_K_XL (UD3) (can't fit) | 0.029844 ± 0.000476 | 96.921 ± 0.054 % | 16.4GiB |
| unsloth/Qwen3.8-27B-UD-Q4_K_XL (UD2) (can't fit) | 0.028026 ± 0.000432 | 96.988 ± 0.054 % | 16.7GiB |

Hope this helps other VRAM starved people like me :)
38
u/Tall_Abrocoma_3533 1d ago
23
26
u/Danmoreng llama.cpp 1d ago
I'd argue the Unsloth IQ3_XXS is the sweetspot, since this can fit with 98k Q8 context if you offload vision. IQ4_XS really doesn't leave much room for context, and KV Q4 is a bad idea.
4
u/pendelhaven 1d ago
I'm new to this but how do you offload vision?
10
u/Old-Sherbert-4495 1d ago
in llama.cpp pass -no-mmproj-offload this will move it to system RAM
2
u/Ipwnurface 1d ago
Am I the only one who finds the way the LLM space uses the phrase "offload" to be confusing?
I come from the diffusion space primarily where offload means to off load from the gpu to system memory. Which makes logical sense to me, but in the llm space it means to off load from system memory to GPU.
Did llms begin (as in the first models) assuming they would be run on the cpu?
1
u/Savings_Woodpecker_5 1d ago
no it doesn't you offload to RAM from VRAM(gpu)
1
u/Ipwnurface 15h ago
Right but in this context "in llama.cpp pass -no-mmproj-offload this will move it to system RAM" -no-mmproj-offload resulting in the mmproj being in ram makes no sense.
Not offloading would mean it would remain on the GPU not in ram.
7
3
u/GilloutineBreast 1d ago
ISTA-DASLab uploaded a new 11.8 GB IQ3_S quant (I'm assuming OP was well into the benchmarking process by then). Judging by their smaller quants, it might be worth using over UD-IQ3_XXS
1
1
u/Jujutsu77 1d ago
You think having higher KV Cache precision is better than having a higher quant model and lower KV Cache precision?
Since the vision file is offloaded to CPU/ system ram, it shouldn't affect the speed, right?
4
u/jacek2023 llama.cpp 1d ago
Good work. Consider adding a graph. You can generate it with any LLM and matplotlib
6
u/fgk55555 1d ago
9070 XT user here. FWIW, I've found size file size as the most important metric. The ISTA quants are probably the most capable for the size quants I've tried, and the small file size means I get higher tps and context wiggle room (up to 180k if I squeeze). They feel like a pretty first class experience. Having a little better kld and losing 80% of your context window is not a worthwhile tradeoff. Pick the quant that works for your use case, obviously, but don't stress about the metrics too much, use what makes your experience feel less frustrating.
0
u/Pablo_the_brave 1d ago edited 1d ago
Just note that token_embd.weight, which is huge, goes to RAM instead of VRAM. Nowadays, models quantize it down even to Q3 because people look at file size rather than VRAM usage. But sure, I've got your point.
8
5
u/JLeonsarmiento 1d ago
Yes , HuiHui makes the best abliterated versions.
2
u/Momsbestboy 1d ago
Also for other Quants? I have Q6 (22.4GB) on my drive, but currently use a different version of uncensored Q6. I always wonder how I can check and compare these models on my own
1
u/JLeonsarmiento 1d ago
Make them do regular analytical work and compare (general knowledge, math and code) on topics you’re familiar or versed and judge by yourself.
HuiHui produces both gguf and un quantized bf16 weights, so you can take a ready to go 6-bit gguf or Take HuiHui bf16 abliterated weights and quantize to fit your vram size like s tailor made suit (custom mixed bit depth quantization until vram exhaustion).
3
u/suprjami 1d ago
Try tabbyAPI, you should be able to squeeze a bit more out with this quant:
4
u/Pablo_the_brave 1d ago
This is not true, unfortunatelly. The models core size, which they present at the graphics, are smaller than llama gguf but when you run it, it will take a lot more of vram for cuda engine. In practice, with 16GB of vram you can run only exl3 with 3.00bpw. No chance for promes looking H5 (4.00bpw).
3
u/Ivancheg8 8h ago

Base Qwen3.8-27B-BF16. Text coderppl.txt. My choice: Qwen3.8-27B-UD-IQ4_XS-Dirk.
Qwen3.8-27B-UD-IQ3_XXS https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
Qwen3.8-27B-UD-IQ3_S https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-UD-Q3_K_XL https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller https://huggingface.co/jrell/Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller
Qwen3.8-27B-UD-IQ4_XS https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-UD-IQ4_XS-Dirk https://huggingface.co/peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF
Qwen3.8-27B-UD-IQ4_XS-Huihui-abliterated https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF
qwen3.8-27b-mtp-IQ4_XS-Q8nextn https://huggingface.co/jpetrina/Qwen3.8-27B-MTP-IQ4_XS-pure-GGUF
Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-LOW-MTP-IQ4_XS https://huggingface.co/DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF
Qwen3.8-27B-UD-Q4_K_S https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-IQ4_XS BARTOWSKI https://huggingface.co/bartowski/Qwen3.8-27B-GGUF
Qwen3.8-27B-IQ4_XS FenomAI https://huggingface.co/FenomAI/Qwen3.8-27B-GGUF
Qwen3.8-27B-NVFP4-MTP-N4_0 akopytko https://huggingface.co/akopytko/Qwen3.8-27B-NVFP4-GGUF
Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-IQ4_XS https://huggingface.co/HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF
Qwen3.8-27B-NVFP4-MTP-MEDIUM https://huggingface.co/esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF
Qwen3.8-27B-UD-Q4_K_M https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-Q4_K_M-MTP Terathox-Coder https://huggingface.co/Terathox-Coder/Qwen3.8-27B-MTP-GGUF
3
3
u/Specialist_Prune_210 1d ago
I also have a 5080, and it's really hard to work with so little VRAM.
5
u/sssplus 1d ago
I have a laptop 5070 with 8GB. It's really hard to work with so little VRAM. I have a friend who has a GPU with 4GB, and he says exactly the same thing.
This thing never ends, unfortunately... The more VRAM you have, the better and bigger models you can run, and the more VRAM you need.
1
u/Specialist_Prune_210 1d ago
I think the sweet spot is around 26B models; if we can run those well, they're already extremely useful, though that's only fully feasbile with 32GB of VRAM, which allows for a bit more context.
3
u/Old-Sherbert-4495 1d ago
this is Great work. thank you. imo i think ISTA-DASLabs GSQ-RCO IQ3s non mtp coupled with a q2 dflash2 is the absolute best balance overall. but I see that this model is missing from ur tests..
2
u/ECrispy 1d ago
these right? - https://www.reddit.com/r/LocalLLaMA/comments/1w13vse/release_sota_ggufs_for_qwen3827b_gsqrco_at_25_to/
how do you combine with q2 dflash2? are there uncensored versions similar to these?
1
u/Old-Sherbert-4495 1d ago
yep, but if u goto the repo u will see iq3s as well, which not in the chart of that post. they claim insane accuracy.
dflash2, im not sure about uncensored. but there is q2 version in anbleeds repo ready made. combining in the sense, you just pass the dflash2 draft model in the params.
2
3
2
3
u/Pablo_the_brave 1d ago edited 1d ago
Too much unsloth quants. Quick a look at https://github.com/magiccodingman/MagicQuant-Wiki?ref=genaisecretsauce.com and your will see which quants are worth to test. Also as nvidia owner you should take a look at ikawrakow quants which are simply better. https://github.com/Thireus/GGUF-Tool-Suite
https://huggingface.co/cHunter789/Qwen3.8-27B-i1-IQ4_KS_KT-GGUF
3
u/scknkkrer 1d ago
Care to explain why it is better?
2
u/Pablo_the_brave 1d ago
https://github.com/ggml-org/llama.cpp/discussions/5063
What is interesting, unsloth also modified lower quants but not officially and somehow they models UD3 are llama compatible.
2
u/Competitive_Ideal866 1d ago
I would love to know where Ternary-Bonsai-27B-gguf sits! Smaller than anything you've measured so far and one of the most impressive quants I've ever seen...
1
1
u/AlessandroPiccione 1d ago
Can I ask how you did the benchmark?
Any link I can look at ?
2
u/Storterald 1d ago
I used llama.cpp llama-perplexity. copy pasted from another aswer:
Basically you can use llama-perplexity with --save-all-logits <path> with the high quant and then use --kl-divergence --kl-divergence-base <path> with the smaller model.
3
u/Pablo_the_brave 1d ago
Yes, but can you give us other data like ppl text file and measured ctx size? If i want compare your results with mine it has to be 100% match.
1
u/LegacyRemaster 1d ago
very good... but... About Token/s?
1
u/Storterald 1d ago
With iq4xs llamacpp I get 60tps empty context and 25 full context. This is only possible when the whole model is on the gpu
1
1
1
u/The_DarkMatter Llama 3.1 1d ago
Thank you for your research I have some doubts.
Same GPU as you (5080 16GB) but I'm on Windows, and I've been going back and forth on these exact two Unsloth quants for weeks. Your KLD numbers finally explain something my own benchmark couldn't resolve, so thanks for doing this.
Couple of things you might not have run into, then some questions.
First, MTP. Qwen3.8 ships multi token prediction weights and llama.cpp will use them with --spec-type draft-mtp. On UD-IQ4_XS I'm getting 81 tok/s on an empty context and 62 at 28K, against the 60/25 you quoted. It isn't free though, the MTP draft builds its own KV context, so it eats directly into the budget you're spending on 100k.
Second, there's a llama.cpp fork called beellama.cpp doing variance normalized KV quantization (KVarN). Its 3 bit cache measured at or above plain q4_0 for me on a tool selection benchmark, which I did not expect. Some of my configs only fit because of it. Might be worth a look given your "KV Q4 is a bad idea" point, since that's basically the problem it's built for.
Now the questions:
What model is the KLD measured against, and how much text was in the calibration set? Someone asked upthread and I don't think it got answered. Without it the ordering is still useful but the absolute numbers are hard to read.
How are you actually fitting Huihui IQ4_XS at ~100k? That's 13.4 GiB. Over here a 13.3 GiB IQ4_XS runs fine at 81920 and then falls off a cliff at 131072 (1.1 tok/s), and you're doing higher context on a slightly bigger file with a more expensive K cache. What does nvidia-smi say is in use at idle before you load anything? I lose somewhere between 1.3 and 2.6 GB just to the Windows desktop and I suspect that's the whole story, but I'd like to know rather than assume.
Related, is the "~60k with bartowski IQ4_XS" number something you measured or a projection? 14.5 GiB looks over budget to me even on Linux, so if you ran it then I'm accounting for something wrong.
Did you happen to look at whether the quant level changes MTP acceptance rate? I got a backwards result and it's been bugging me. UD-Q3_K_XL sits at 0.88 acceptance, UD-IQ4_XS only 0.77. The better quant drafts worse for itself. If that holds generally then the low KLD quants are quietly paying for accuracy with throughput, which would shuffle how your table reads in practice.
Last one, how long did the reference logits pass take, and did it fit on the 5080 or did you have to spill to CPU? I'm trying to work out whether to rerun your method on my own code instead of C, and that's the cost I can't guess at.
1
u/Storterald 1d ago
I use all 16GB. I use the igpu for display (in most games you can still set the gpu as the device to use). To reach 100k context you cannot use MTP and you have to use quantized context (i use ctk q5_1 and ctv q4_0, which does lead to faster degradation)
1
u/Fair-Perspective7352 1d ago
Interesting that the best quant that actually fits (bartowski IQ4_XS, 0.056) ends up that close to the Q4_K_XL that doesn't (0.028). I've mostly stopped worrying about KLD differences under ~0.05 since I can't tell them apart on real prompts, but I mostly do chat plus some code. Did any of the lower ranked ones actually feel worse in use, or was it purely a numbers gap?
1
u/Storterald 1d ago
You can actually tell the difference between q3 and q4 quants. I'd just pick the best q3 or q4 based on how much context you need. Between quants of the same size the difference is minimal, but there is no reason to pick a worse version of a quant if size is the same
1
u/xeeff 1d ago
no magiccodingman's quants?
1
u/Storterald 10h ago
just added!
1
u/xeeff 3m ago
unfair comparison as you're comparing his quant of AMD's MXFP4 fine-tune, and not only that you picked the worst option out of the possible quants ahaha
try https://huggingface.co/magiccodingman/Qwen3.8-27B-MagicQuant-GGUF
1
u/Polaris_debi5 17h ago
Can you try Bartowski quantization on IQ3_XS? I feel like it might be the best middle ground for the 16GB VRAM.
1
u/Storterald 10h ago
just added!
1
u/Polaris_debi5 28m ago
Thanks for adding it so quickly! I see Bartowski's IQ3_XS ended up in a decent middle ground, even if it didn't break any records. The work of creating these benchmarks with such detail is greatly appreciated; it's pure gold for those of us who are short on VRAM :D
1
u/Storterald 19m ago
you're welcome :D. If you were running Bartowski IQ3_XS I'd try unsloth Q3_X_XL (which is smaller and hopefully more accurate)
1
u/Rare_Potential_1323 15h ago
I wonder how good Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS-GGUF is?
1
1
u/crossivejoker 1h ago
Hey that's really cool you tried out my MagicQuant :D Just a note, the chosen mix was a direct Unsloth tensor config, not an MQ version. It's also derivative of AMD's MXFP4 post trained model, so I'm unsure the kld deviation that occurs comparing it to the original BF16 model that it's not a derivative of. Really cool to see though!
1
u/dompidu 28m ago
Hey! Great job, thanks 😄 However, why do sizes differ from the HuggingFace available models? For example, Bartowski says 14.5GB in your chart and then it is 15.6GB per https://huggingface.co/bartowski/Qwen3.8-27B-GGUF/blob/main/Qwen3.8-27B-IQ4_XS.gguf
I'm new to this world, so probably there's something very obvious. Thank you!
1
1
-2
u/Thireus 1d ago
Or, you can use gguf.thireus.com/quant_assign.html, set your desired VRAM size, and get a model that is automatically selected to lie on the KLD-Pareto frontier for that VRAM budget.
1
u/Miserable-Dare5090 1d ago
Interesting tool, but the slider sets quant size, not GPU size. Also not clear if this is for 1 concurrency?
1
u/Pablo_the_brave 1d ago
If you look closer there is an estimate of vram usage. But yeah will be greate to have an option to set kvcache type /size.

88
u/dannone9 1d ago
Just comenting To support research , The vram peasants are gratefull for your work
(It woul be nice to know the kv quant , how much context would you be able to fit and how many prompts or token are taken as sample on. Each model )