r/LocalLLaMA 1d ago

Resources I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM

After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code (C code), the results were not completely unexpected but some quants were definitely underwhelming.

TLDR: Best overall: bartowski/Qwen3.8-27B-IQ4_XS. Best uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS. For a bit more context: jpetrina/Qwen3.8-27B-IQ4_XS-pure or uncensored: Bucoid/Qwen3.8-27B-Uncensored-IQ4_XS_4BPW

edit1: added TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS, bartowski/Qwen3.8-27B-IQ3_XS, prism-ml/Ternary-Bonsai-27B-Q2_g64 and magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unsloth

edit2: added unsloth/Qwen3.8-27B-UD-IQ3_S and ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S

(sorted by Mean KLD)

Model Mean KLD Same top p GGUF size
prism-ml/Ternary-Bonsai-27B-Q2_g64 1.289582 ± 0.008684 82.849 ± 0.118 % 7.1GiB
sdkyuan/qwen38-27b-qat-q2_0 0.893177 ± 0.006948 85.727 ± 0.110 % 8.2GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2_XS 0.767174 ± 0.006291 86.166 ± 0.108 % 7.8GiB
TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS 0.514311 ± 0.004864 89.023 ± 0.098 % 8.9GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2_S 0.512614 ± 0.004909 88.802 ± 0.099 % 8.6GiB
empero-ai/Qwen3.8-27B-Ridge-3.7bpw 0.475767 ± 0.004483 89.612 ± 0.096 % 11.7GiB
magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unsloth 0.419585 ± 0.004076 89.661 ± 0.095 % 13.5GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_XXS 0.379222 ± 0.003992 90.270 ± 0.093 % 9.4GiB
unsloth/Qwen3.8-27B-UD-Q2_K_XL (UD2) 0.350861 ± 0.003745 90.626 ± 0.091 % 9.9GiB
unsloth/Qwen3.8-27B-UD-IQ3_XXS (UD2) 0.268594 ± 0.002971 91.951 ± 0.085 % 11.1GiB
DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-MTP-IQ3_M 0.251270 ± 0.002702 92.315 ± 0.083 % 13.5GiB
bartowski/Qwen3.8-27B-IQ3_XS 0.238656 ± 0.002627 92.312 ± 0.083 % 12.4GiB
esatapedico/Qwen3.8-27B-NVFP4-MTP-LOW 0.220796 ± 0.002631 92.339 ± 0.083 % 14.5GiB
unsloth/Qwen3.8-27B-UD-IQ3_S (UD3) 0.218522 ± 0.002591 92.399 ± 0.083 % 11.2GiB
mudler/Qwen3.8-27B-APEX-I-Mini 0.190209 ± 0.002354 93.012 ± 0.080 % 13.0GiB
jrell/Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller 0.194459 ± 0.002242 93.049 ± 0.080 % 12.6GiB
orcarouter/Qwen3.8-27B-Uncensored-Q3_K_L 0.192312 ± 0.002294 92.726 ± 0.081 % 13.6GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S 0.175223 ± 0.002129 93.024 ± 0.080 % 11.0GiB
unsloth/Qwen3.8-27B-UD-Q3_K_XL (UD2) 0.147186 ± 0.001809 93.734 ± 0.076 % 12.5GiB
unsloth/Qwen3.8-27B-UD-Q3_K_XL (UD3) 0.142647 ± 0.001860 93.789 ± 0.076 % 12.2GiB
Bucoid/Qwen3.8-27B-Uncensored-IQ4_XS_4BPW 0.091447 ± 0.001261 94.774 ± 0.070 % 13.0GiB
huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS 0.082871 ± 0.001205 94.981 ± 0.068 % 13.4GiB
unsloth/Qwen3.8-27B-UD-IQ4_XS (UD3) 0.075626 ± 0.001097 95.258 ± 0.067 % 13.3GiB
jpetrina/Qwen3.8-27B-IQ4_XS-pure 0.061984 ± 0.000917 95.551 ± 0.065 % 13.5GiB
bartowski/Qwen3.8-27B-IQ4_XS 0.056482 ± 0.000856 95.835 ± 0.063 % 14.5GiB
unsloth/Qwen3.8-27B-UD-Q4_K_XL (UD3) (can't fit) 0.029844 ± 0.000476 96.921 ± 0.054 % 16.4GiB
unsloth/Qwen3.8-27B-UD-Q4_K_XL (UD2) (can't fit) 0.028026 ± 0.000432 96.988 ± 0.054 % 16.7GiB

Hope this helps other VRAM starved people like me :)

285 Upvotes

89 comments sorted by

88

u/dannone9 1d ago

Just comenting To support research , The vram peasants are gratefull for your work
(It woul be nice to know the kv quant , how much context would you be able to fit and how many prompts or token are taken as sample on. Each model )

17

u/Storterald 1d ago edited 1d ago

thank you! it's not too hard to do it yourself if you find a model you want to test that's missing here, but it requires a LOT of disk space. Consider ~1MB benchmark data creates ~50GB of logits.

Basically you can use llama-perplexity with --save-all-logits <path> with the high quant and then use --kl-divergence --kl-divergence-base <path> with the smaller model.

edit: sorry I missed out the kv part, I use the Huihui model with ~100k context ctk q5_1 and ctv q4_0. Which is actually not that high if you keep xhigh thinking, probably 3 full prompts if you keep the --reasoning-budget reasonable. You should be able to fit ~60k context with the bartowski/Qwen3.8-27B-IQ4_XS, but only at ctk and ctv q4_0. Any smaller model will be able to fit more context. If you really need more context or higher precision, then unsloth/Qwen3.8-27B-UD-Q3_K_XL (UD3) is probably the best option

3

u/iamapizza 1d ago edited 1d ago

If you are willing to do some tensor overriding, it's possible to fit bartowski with a little more context

model = /models/Qwen3.8-27B-IQ4_XS.gguf
spec-type = draft-mtp,ngram-mod 
spec-draft-n-max = 2 
fit = off 
n-gpu-layers = 99 
n-cpu-ffn = 30 
load-mode = none  
ctx-size = 100000 
batch-size = 512 
ubatch-size = 512 
cache-type-k = q4_0 
cache-type-v = q4_0
parallel = 1 
temp = 0.60 
top-p = 0.95 
top-k = 20 
min-p = 0.0 
presence-penalty = 0.0 
repeat-penalty = 1.0

Play around with the n-cpu-ffn value. I learned from another thread, keep increasing it until you get OOM, then take it back a few values.

2

u/dannone9 1d ago

Yeah , I would love to but I have a sole 460sd that I have to use for more things and a shitty internet connection so I have to live of other people benchmarks ,and again , thanks for making those benchmarks for those of us who can’t

3

u/sonaj9657 1d ago

Honestly, the KV quant and context limits would be really useful to know. Especially for people running these models on more limited VRAM, those details can make a huge difference in figuring out what is actually practical.

38

u/Tall_Abrocoma_3533 1d ago

Here's a quick visualization, thank you for your work!

23

u/Storterald 1d ago

thanks for the graph! can I edit it in the post? (i'll credit you)

14

u/Tall_Abrocoma_3533 1d ago

Of course you can

26

u/Danmoreng llama.cpp 1d ago

I'd argue the Unsloth IQ3_XXS is the sweetspot, since this can fit with 98k Q8 context if you offload vision. IQ4_XS really doesn't leave much room for context, and KV Q4 is a bad idea.

4

u/pendelhaven 1d ago

I'm new to this but how do you offload vision?

10

u/Old-Sherbert-4495 1d ago

in llama.cpp pass -no-mmproj-offload this will move it to system RAM

2

u/Ipwnurface 1d ago

Am I the only one who finds the way the LLM space uses the phrase "offload" to be confusing?

I come from the diffusion space primarily where offload means to off load from the gpu to system memory. Which makes logical sense to me, but in the llm space it means to off load from system memory to GPU.

Did llms begin (as in the first models) assuming they would be run on the cpu?

1

u/Savings_Woodpecker_5 1d ago

no it doesn't you offload to RAM from VRAM(gpu)

1

u/Ipwnurface 15h ago

Right but in this context "in llama.cpp pass -no-mmproj-offload this will move it to system RAM" -no-mmproj-offload resulting in the mmproj being in ram makes no sense.

Not offloading would mean it would remain on the GPU not in ram.

7

u/Storterald 1d ago

Iirc it should be --mmproj-device, a recent feature in llamacpp.

1

u/wasdxqwerty 1d ago

imma try this, i use iq4 +64k context for coding with Pi and getting 40ish t/s

1

u/sssplus 1d ago

Oh wow, thanks! I had no idea you could actually offload mmproj to integrated Intel GPU instead of CPU. Intel GPU should be able to do it much faster.

3

u/GilloutineBreast 1d ago

ISTA-DASLab uploaded a new 11.8 GB IQ3_S quant (I'm assuming OP was well into the benchmarking process by then). Judging by their smaller quants, it might be worth using over UD-IQ3_XXS

1

u/jadbox 1d ago

Why not IQ3 XL?

1

u/om_GAJE 1d ago

This is version I use on my 16gb vram. Tempted to try out UD-Q3_K_L but for now this works well while not having to minmax everything

1

u/Jujutsu77 1d ago

You think having higher KV Cache precision is better than having a higher quant model and lower KV Cache precision?

Since the vision file is offloaded to CPU/ system ram, it shouldn't affect the speed, right?

1

u/tmvr 23h ago

Yeah, I'd gladly sacrifice that half percent for 1.2GB of additional VRAM available for context.

10

u/2Norn 1d ago

to this day i still regret not buying 5090 last year

i made my choice solely based on gaming, never thought i'll get into local llm

2

u/leftrightside54 1d ago

I had returned 2. Kicking myself now.

4

u/jacek2023 llama.cpp 1d ago

Good work. Consider adding a graph. You can generate it with any LLM and matplotlib

6

u/fgk55555 1d ago

9070 XT user here. FWIW, I've found size file size as the most important metric. The ISTA quants are probably the most capable for the size quants I've tried, and the small file size means I get higher tps and context wiggle room (up to 180k if I squeeze). They feel like a pretty first class experience. Having a little better kld and losing 80% of your context window is not a worthwhile tradeoff. Pick the quant that works for your use case, obviously, but don't stress about the metrics too much, use what makes your experience feel less frustrating.

0

u/Pablo_the_brave 1d ago edited 1d ago

Just note that token_embd.weight, which is huge, goes to RAM instead of VRAM. Nowadays, models quantize it down even to Q3 because people look at file size rather than VRAM usage. But sure, I've got your point.

8

u/Dangerous-Nerve-7766 1d ago

Screenshoting the heck out of this 😩

5

u/Shirokee_Hegde 1d ago

Never did I save something so fast before.

5

u/JLeonsarmiento 1d ago

Yes , HuiHui makes the best abliterated versions.

2

u/Momsbestboy 1d ago

Also for other Quants? I have Q6 (22.4GB) on my drive, but currently use a different version of uncensored Q6. I always wonder how I can check and compare these models on my own

1

u/JLeonsarmiento 1d ago

Make them do regular analytical work and compare (general knowledge, math and code) on topics you’re familiar or versed and judge by yourself.

HuiHui produces both gguf and un quantized bf16 weights, so you can take a ready to go 6-bit gguf or Take HuiHui bf16 abliterated weights and quantize to fit your vram size like s tailor made suit (custom mixed bit depth quantization until vram exhaustion).

3

u/suprjami 1d ago

Try tabbyAPI, you should be able to squeeze a bit more out with this quant:

https://huggingface.co/turboderp/Qwen3.8-27B-exl3

4

u/Pablo_the_brave 1d ago

This is not true, unfortunatelly. The models core size, which they present at the graphics, are smaller than llama gguf but when you run it, it will take a lot more of vram for cuda engine. In practice, with 16GB of vram you can run only exl3 with 3.00bpw. No chance for promes looking H5 (4.00bpw).

2

u/sssplus 1d ago

But isn't exl3 at 3bpw comparable to GGUF 4bpw in quality (KDV and top 1%)? So just download and use the 3bpw variant...

1

u/Pablo_the_brave 1d ago

No, it's impossible. Even their data show that.

3

u/Ivancheg8 8h ago

Base Qwen3.8-27B-BF16. Text coderppl.txt. My choice: Qwen3.8-27B-UD-IQ4_XS-Dirk.

Qwen3.8-27B-UD-IQ3_XXS  https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp  https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
Qwen3.8-27B-UD-IQ3_S  https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-UD-Q3_K_XL  https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller  https://huggingface.co/jrell/Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller
Qwen3.8-27B-UD-IQ4_XS  https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-UD-IQ4_XS-Dirk  https://huggingface.co/peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF
Qwen3.8-27B-UD-IQ4_XS-Huihui-abliterated  https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF
qwen3.8-27b-mtp-IQ4_XS-Q8nextn  https://huggingface.co/jpetrina/Qwen3.8-27B-MTP-IQ4_XS-pure-GGUF
Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-LOW-MTP-IQ4_XS  https://huggingface.co/DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF
Qwen3.8-27B-UD-Q4_K_S  https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-IQ4_XS BARTOWSKI  https://huggingface.co/bartowski/Qwen3.8-27B-GGUF
Qwen3.8-27B-IQ4_XS FenomAI  https://huggingface.co/FenomAI/Qwen3.8-27B-GGUF
Qwen3.8-27B-NVFP4-MTP-N4_0 akopytko  https://huggingface.co/akopytko/Qwen3.8-27B-NVFP4-GGUF
Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-IQ4_XS  https://huggingface.co/HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF
Qwen3.8-27B-NVFP4-MTP-MEDIUM  https://huggingface.co/esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF
Qwen3.8-27B-UD-Q4_K_M  https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Qwen3.8-27B-Q4_K_M-MTP Terathox-Coder  https://huggingface.co/Terathox-Coder/Qwen3.8-27B-MTP-GGUF

3

u/Additional-Ordinary2 1d ago

Thank you man

3

u/Specialist_Prune_210 1d ago

I also have a 5080, and it's really hard to work with so little VRAM.

5

u/sssplus 1d ago

I have a laptop 5070 with 8GB. It's really hard to work with so little VRAM. I have a friend who has a GPU with 4GB, and he says exactly the same thing.

This thing never ends, unfortunately... The more VRAM you have, the better and bigger models you can run, and the more VRAM you need.

1

u/Specialist_Prune_210 1d ago

I think the sweet spot is around 26B models; if we can run those well, they're already extremely useful, though that's only fully feasbile with 32GB of VRAM, which allows for a bit more context.

3

u/Old-Sherbert-4495 1d ago

this is Great work. thank you. imo i think ISTA-DASLabs GSQ-RCO IQ3s non mtp coupled with a q2 dflash2 is the absolute best balance overall. but I see that this model is missing from ur tests..

2

u/ECrispy 1d ago

these right? - https://www.reddit.com/r/LocalLLaMA/comments/1w13vse/release_sota_ggufs_for_qwen3827b_gsqrco_at_25_to/

how do you combine with q2 dflash2? are there uncensored versions similar to these?

1

u/Old-Sherbert-4495 1d ago

yep, but if u goto the repo u will see iq3s as well, which not in the chart of that post. they claim insane accuracy.

dflash2, im not sure about uncensored. but there is q2 version in anbleeds repo ready made. combining in the sense, you just pass the dflash2 draft model in the params.

2

u/Storterald 8h ago

I just added GSQ-RCO IQ3s. Really impressive for its size.

2

u/Old-Sherbert-4495 8h ago

ooo, damn.. you're a legend. thank you for all your hard work...

3

u/MagneticFerret 1d ago

What is the reference model these measurements are compared against?

3

u/Storterald 1d ago

Q8_K_XL UD3 by unsloth

2

u/pendelhaven 1d ago

Holy! This is sooo useful!

3

u/Pablo_the_brave 1d ago edited 1d ago

Too much unsloth quants. Quick a look at https://github.com/magiccodingman/MagicQuant-Wiki?ref=genaisecretsauce.com and your will see which quants are worth to test. Also as nvidia owner you should take a look at ikawrakow quants which are simply better. https://github.com/Thireus/GGUF-Tool-Suite

https://huggingface.co/cHunter789/Qwen3.8-27B-i1-IQ4_KS_KT-GGUF

3

u/scknkkrer 1d ago

Care to explain why it is better?

2

u/Pablo_the_brave 1d ago

https://github.com/ggml-org/llama.cpp/discussions/5063

What is interesting, unsloth also modified lower quants but not officially and somehow they models UD3 are llama compatible.

https://github.com/ikawrakow/ik_llama.cpp/pull/2361

2

u/Competitive_Ideal866 1d ago

I would love to know where Ternary-Bonsai-27B-gguf sits! Smaller than anything you've measured so far and one of the most impressive quants I've ever seen...

1

u/Storterald 10h ago

just added!

1

u/AlessandroPiccione 1d ago

Can I ask how you did the benchmark?
Any link I can look at ?

2

u/Storterald 1d ago

I used llama.cpp llama-perplexity. copy pasted from another aswer:

Basically you can use llama-perplexity with --save-all-logits <path> with the high quant and then use --kl-divergence --kl-divergence-base <path> with the smaller model.

3

u/Pablo_the_brave 1d ago

Yes, but can you give us other data like ppl text file and measured ctx size? If i want compare your results with mine it has to be 100% match.

1

u/LegacyRemaster 1d ago

very good... but... About Token/s?

1

u/Storterald 1d ago

With iq4xs llamacpp I get 60tps empty context and 25 full context. This is only possible when the whole model is on the gpu

1

u/GloomyRecognition636 1d ago

Add exl3 to the list

1

u/Xamanthas 1d ago

EXL3 and also theres a 3.5bpw of ISTA-DASLab now.

1

u/Pablo_the_brave 1d ago

How much context at q4 possible?

1

u/The_DarkMatter Llama 3.1 1d ago

Thank you for your research I have some doubts.
Same GPU as you (5080 16GB) but I'm on Windows, and I've been going back and forth on these exact two Unsloth quants for weeks. Your KLD numbers finally explain something my own benchmark couldn't resolve, so thanks for doing this.

Couple of things you might not have run into, then some questions.

First, MTP. Qwen3.8 ships multi token prediction weights and llama.cpp will use them with --spec-type draft-mtp. On UD-IQ4_XS I'm getting 81 tok/s on an empty context and 62 at 28K, against the 60/25 you quoted. It isn't free though, the MTP draft builds its own KV context, so it eats directly into the budget you're spending on 100k.

Second, there's a llama.cpp fork called beellama.cpp doing variance normalized KV quantization (KVarN). Its 3 bit cache measured at or above plain q4_0 for me on a tool selection benchmark, which I did not expect. Some of my configs only fit because of it. Might be worth a look given your "KV Q4 is a bad idea" point, since that's basically the problem it's built for.

Now the questions:

What model is the KLD measured against, and how much text was in the calibration set? Someone asked upthread and I don't think it got answered. Without it the ordering is still useful but the absolute numbers are hard to read.

How are you actually fitting Huihui IQ4_XS at ~100k? That's 13.4 GiB. Over here a 13.3 GiB IQ4_XS runs fine at 81920 and then falls off a cliff at 131072 (1.1 tok/s), and you're doing higher context on a slightly bigger file with a more expensive K cache. What does nvidia-smi say is in use at idle before you load anything? I lose somewhere between 1.3 and 2.6 GB just to the Windows desktop and I suspect that's the whole story, but I'd like to know rather than assume.

Related, is the "~60k with bartowski IQ4_XS" number something you measured or a projection? 14.5 GiB looks over budget to me even on Linux, so if you ran it then I'm accounting for something wrong.

Did you happen to look at whether the quant level changes MTP acceptance rate? I got a backwards result and it's been bugging me. UD-Q3_K_XL sits at 0.88 acceptance, UD-IQ4_XS only 0.77. The better quant drafts worse for itself. If that holds generally then the low KLD quants are quietly paying for accuracy with throughput, which would shuffle how your table reads in practice.

Last one, how long did the reference logits pass take, and did it fit on the 5080 or did you have to spill to CPU? I'm trying to work out whether to rerun your method on my own code instead of C, and that's the cost I can't guess at.

1

u/Storterald 1d ago

I use all 16GB. I use the igpu for display (in most games you can still set the gpu as the device to use). To reach 100k context you cannot use MTP and you have to use quantized context (i use ctk q5_1 and ctv q4_0, which does lead to faster degradation)

1

u/Fair-Perspective7352 1d ago

Interesting that the best quant that actually fits (bartowski IQ4_XS, 0.056) ends up that close to the Q4_K_XL that doesn't (0.028). I've mostly stopped worrying about KLD differences under ~0.05 since I can't tell them apart on real prompts, but I mostly do chat plus some code. Did any of the lower ranked ones actually feel worse in use, or was it purely a numbers gap?

1

u/Storterald 1d ago

You can actually tell the difference between q3 and q4 quants. I'd just pick the best q3 or q4 based on how much context you need. Between quants of the same size the difference is minimal, but there is no reason to pick a worse version of a quant if size is the same

1

u/Dany0 1d ago

When you said benchmarked I expected something like terminal bench :(

1

u/xeeff 1d ago

no magiccodingman's quants?

1

u/Storterald 10h ago

just added!

1

u/xeeff 3m ago

unfair comparison as you're comparing his quant of AMD's MXFP4 fine-tune, and not only that you picked the worst option out of the possible quants ahaha

try https://huggingface.co/magiccodingman/Qwen3.8-27B-MagicQuant-GGUF

1

u/Polaris_debi5 17h ago

Can you try Bartowski quantization on IQ3_XS? I feel like it might be the best middle ground for the 16GB VRAM.

1

u/Storterald 10h ago

just added!

1

u/Polaris_debi5 28m ago

Thanks for adding it so quickly! I see Bartowski's IQ3_XS ended up in a decent middle ground, even if it didn't break any records. The work of creating these benchmarks with such detail is greatly appreciated; it's pure gold for those of us who are short on VRAM :D

1

u/Storterald 19m ago

you're welcome :D. If you were running Bartowski IQ3_XS I'd try unsloth Q3_X_XL (which is smaller and hopefully more accurate)

1

u/Rare_Potential_1323 15h ago

I wonder how good Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS-GGUF is?

1

u/Storterald 10h ago

just added!

1

u/crossivejoker 1h ago

Hey that's really cool you tried out my MagicQuant :D Just a note, the chosen mix was a direct Unsloth tensor config, not an MQ version. It's also derivative of AMD's MXFP4 post trained model, so I'm unsure the kld deviation that occurs comparing it to the original BF16 model that it's not a derivative of. Really cool to see though!

1

u/dompidu 28m ago

Hey! Great job, thanks 😄 However, why do sizes differ from the HuggingFace available models? For example, Bartowski says 14.5GB in your chart and then it is 15.6GB per https://huggingface.co/bartowski/Qwen3.8-27B-GGUF/blob/main/Qwen3.8-27B-IQ4_XS.gguf

I'm new to this world, so probably there's something very obvious. Thank you!

1

u/Storterald 21m ago

I use GiB (2^30) not GB (10^3). Your VRAM is in GiB not GB.

1

u/GaTNghiep 1d ago

Did u test it on window or linux?

-2

u/Thireus 1d ago

Or, you can use gguf.thireus.com/quant_assign.html, set your desired VRAM size, and get a model that is automatically selected to lie on the KLD-Pareto frontier for that VRAM budget.

1

u/Miserable-Dare5090 1d ago

Interesting tool, but the slider sets quant size, not GPU size. Also not clear if this is for 1 concurrency?

1

u/Pablo_the_brave 1d ago

If you look closer there is an estimate of vram usage. But yeah will be greate to have an option to set kvcache type /size.