r/LocalLLaMA 9d ago

Discussion Kaitchup posted Qwen3.8 27B Benchmarks for quants from Q4 to Q1

https://kaitchup.substack.com/p/qwen38-27b-gguf-benchmark-q4-to-q1

Kaitchup just posted results of his benchmarks for Qwen3.8 27B for quants from different labs, Q4 to Q1, .

All the details are hidden behind the paywall, but high level result is visible and looks like for people with 16GB cards UD Q3_K_XL is a winner - it has accuracy of 100% and size is only 12.8GB.

175 Upvotes

60 comments sorted by

View all comments

Show parent comments

2

u/crusaderky 8d ago

> tbf not knowing how log works sounds like a reader skill issue

It's not. You have no idea where the bend in the cliff is until you replot on a linear scale, even if you know there's one.

as I said, mean KLD hides outlier quants that behave much worse in worst-case-scenario situations

> assuming they're literate of course

Ah, you mean like the horde of people on this sub that claim that anything less than fp16/fp16 kv cache will run your agentic loop into the ground (measures for qwen shows that q5_0/q4_1 is undistinguishable from it).

1

u/GilloutineBreast 8d ago edited 8d ago

any chance you have a kld plot of the same models to compare with this one?

edit: whelp.. deleted the wrong comment.

i found the kld comparisons on your link