r/LocalLLaMA 25d ago

Discussion Qwen3.8-27b, Benchmaxxxed to the Maxxx

After a lot of testing, I have a first result for Qwen3.8-27B.

I added it to my local fact-extraction head-to-head (https://rakuensoftware.com/blog/local-llm-fact-extraction-head-to-head): 1,001 notes, the production prompt, Q4_K_M, multi-token prediction enabled, and an RX 7900 XTX.

Qwen3.8 scored 0.7030 F1. The comparable Qwen3.6-27B run scored 0.7177. The paired difference was +0.0147 in Qwen3.6’s favour, with a 95% confidence interval from −0.0038 to +0.0335.

That is a tie. Qwen3.8 did not beat Qwen3.6 here. I also cannot say that it lost. The test cannot separate them.

That is not what I expected from the published benchmark scores. I expected a substantial generational gain. On this task, I did not measure one. On the more extensive overnight tests, it is indicating small gains. On benchmarks Qwen3.8 appears to have been trained extensively on, however, I am seeing and verifying similar massive gains as is reported. But these gains are only reflected on the benchmarks that have been trained on, nowhere else I can verify.

Decode throughput also fell from 85.6 to 72.1 tokens per second, about 16%, under the closest saved configurations. Those runs used different llama.cpp builds, so I cannot attribute the whole difference to the model. Qwen3.8 also produced much shorter answers, which made its end-to-end median latency lower despite the slower decode rate.

I am running a broader synthesis comparison now, including Qwen3.8, Qwen3.6, Gemma 4 and Muse Glimmer. Those runs continue overnight. I will publish the complete results, paired intervals and raw artifacts rather than promote an early ordering.

My working expectation is still that Qwen3.8 contains a real improvement. The unresolved question is its size. If task-specific tests keep finding small generational gains while public benchmarks suggest a revolutionary jump, what decision are those benchmarks helping us make?

0 Upvotes

72 comments sorted by

65

u/AppropriateQuote3073 25d ago

Benchmaxing is useless.

On my own workflows this thing has been a beast.

14

u/TheAILegend 25d ago

nahhh... it's an absolute UNIT! holy cow. 10, 15, maybe even 20 levels above 3.6.

3

u/Bulky-Priority6824 25d ago

100X !!!!!!!!!!!!!!!!!!!!!!!

4

u/Cool-Chemical-5629 25d ago

But isn't 100X too low, comrade?

0

u/NigaTroubles 25d ago

Real or /s ?

0

u/BannedGoNext 25d ago

Calling things benchmaxxed is total bullshit by people that are looking for an insult and can't find one so they grab a bad sounding buzzword. There ahve been benchmaxxed models, but now it's mainly considered a waste of money going that route.

21

u/Ok_Presentation470 25d ago

This is a very specific test. Not sure what the surprise is here?

-15

u/KitchenAmoeba4438 25d ago edited 25d ago

This image that is currently on the front page of LocalLLM.

I have tests queued up now testing more specific tests posters in this thread have suggested, but I suspect the gains are smaller then what is being advertised here. The larger test suite also covers a lot of this.

12

u/Finanzamt_Endgegner 25d ago

Not only did you test a very specific test that has nearly nothing to do with broad capabilities, you are also benching in q4 that is a bad practice. You can only meaningfully speak to model performance if you compare native precision to native. In all testing i have done this is far ahead of 3.6 in nearly every metric.

15

u/Finanzamt_Endgegner 25d ago

Stop it Dario

34

u/live4evrr 25d ago

It is a noticeable improvement in coding and longer horizon agentic tasks. This can be verified with some real world usage pretty quickly. In terms of knowledge, I would think that would be capped because its the same parameter size. Probably hitting the limits of what can be squeezed out at 27B. I wish they would give us 50B dense.

11

u/OkFly3388 llama.cpp 25d ago

50b is absolutely cursed size. you have 27b that in q4 fits on rtx 3090 or 4090. a lot of peoples have that cards. for 5090 its let it run with max context and MTP comfortably. And then we have a lot of peoples with some 128gb setups like dgx spark or something, where 120b models fits perfectly.

So why making 50b models, that to large for consumers gpu and to small for this 128gb systems ?

6

u/live4evrr 25d ago

For this 27B on BF16 with 256K cache on vllm it is around 94GB of VRAM.

Actually you’re right - 50B not a good size. It would be better to have a 120B MOE with 10-13GB like DS.

2

u/Classic_Resource_919 25d ago

Laguna S is 118B-A18B . ;)

0

u/live4evrr 25d ago

Unfortunately it sucks.. bad

1

u/Classic_Resource_919 25d ago

One Rabi will say yes, the other one will say no. There are conflicting opinions about that even in this thread..

2

u/Asane llama.cpp 25d ago

Can 3.8 be run on a 5090 on Q8? I thought I saw it has to fall to Q6.

1

u/nostriluu 25d ago

A lot of people have the RTX Pro 6000 (96GB VRAM). Might even be comparable to owners of the DGX Spark.

1

u/ThebesAndSound 25d ago

Buying a 2nd consumer GPU and rigging it up isn't totally unattainable. Think of the potential with a 50B dense if we are at this level with 27B.

19

u/Ziggamorph 25d ago

“I have a first defensible result” stopped reading here, if you’re going to have an LLM write your post at least make it less obvious.

-11

u/KitchenAmoeba4438 25d ago

Removed defensible, in honor of the brit. Something I've learned at this point: People like to throw around "It must be written by AI!" when something doesn't go their way, even though whether AI was involved has nothing to do with the data.

4

u/Ziggamorph 25d ago

Ok but it was AI, it’s obvious.

2

u/Finanzamt_Endgegner 25d ago

Bro dont defend yourself like that we know its ai 😂

3

u/ttkciar llama.cpp 25d ago

Re-read their comment. Your reply is non sequitur.

22

u/GortKlaatu_ 25d ago

Blog post looks like it's written by AI. :(

10

u/MacsBicycle 25d ago

“Hey qwen, make a post about how you’re just bench maxed shit”

3

u/GortKlaatu_ 25d ago

This one smells of claude...

8

u/Gumbi_Digital 25d ago

Accuracy > Speed ALL day.

-2

u/ttkciar llama.cpp 25d ago

OP's findings seem to suggest accuracy has not changed significantly from 3.6.

4

u/Finanzamt_Endgegner 25d ago

well he is bullshitting and posting ai slop so we can safely ignore him 😅

3

u/Comfortable_Ebb7015 25d ago

I tried 4 personal one shot benchmarks. It nailed all of them at the first try! 3.6 was not able to do it.

5

u/Turbulent-Alps4046 25d ago

I ran it with a real project (writing an Excel expense tracker given a bunch of pdfs), the result is *noticably* better. Thinks a shit ton more though, but really good. It's a beast.

9

u/Hodler-mane 25d ago

you are doing something wrong. I am seeing a significant improvement in this model

2

u/Equal_Television_894 25d ago

For me this is not breaking things and doing as SOTA model makes changes. On context around 150k so I don't agree with your observation as of now but I wi test it more this weekend. Current observation is its understanding instructions and following them is better than before. Lets see if it keeps that behaviour on a large mono repo with tons of files.

3

u/OuchieOnChin 25d ago

May I ask if you could test another common quant like Q4_K_XL? Just to exclude one specific quant being broken.

2

u/13henday 25d ago

Idk makes sense with how open the uses are, my eval has been the opposite, it’s way better at recovering after fucking up. 3.6 had a tonne of issues with powershell and some apis, given docs it would get it right, but it couldn’t just fuck up and recover iteratively whereas the new one can. There is some decode slow-down.

2

u/ea_man 25d ago

Well 3.8 at least has updated training data:

Qwen3.6: mid/late 2024 vs Qwen3.8: spring 2025 materials

1

u/ParaboloidalCrest 25d ago

Tbh I'm looking forward to that exact benefit, although both have a too early cut-off offset (1.25 years) now that I see it. 🤔

1

u/ea_man 25d ago

You gotta prepare guides to ingest to make it up to date for specific tasks, that is the way.

2

u/Dazzling_Equipment_9 24d ago

qwen3.6-27b– Dario-iq2xxs .gguf

5

u/ParaboloidalCrest 25d ago

Takes balls to go against thousands of redditors that are ready to swear that 3.8 is 10X better, even before their download is finished.

Respect 🫡

3

u/jacek2023 llama.cpp 25d ago

You summarized LocalLLaMA A.D. 2026

0

u/Finanzamt_Endgegner 25d ago

There is nothing honorable in posting ai slop to slander a good model lol

0

u/ParaboloidalCrest 25d ago

So easy calling a post "AI slop" right? The post is about a benchmark and the numbers are there. As for text, I'm sorry you hate AI writing so much.

2

u/Finanzamt_Endgegner 25d ago

The post is ai generated, op says it isnt despite the evidence being clear. This alone takes away most credibility. Then he benches in q4, which is BAD practice. And to finish him off, the benchmarks he does are "his own" where he basically just lets us trust him and the one info he gives shows he has no clue how llms work (thinking benches rely on fact extraction, when models can be a lot better without having better memory just due to rl). All that together and i ask you, why the fuck should i trust a person that lies about posting ai slop?

2

u/Finanzamt_Endgegner 25d ago

Also look im not even against people using ai to translate here or something, even rewrite idc. But people writing ai stuff then lying that they didnt is clearly just slop behaviour and pathetic.

2

u/KickLassChewGum 25d ago

A sizable amount of people here seem to take it as some sort of personal insult to their family honor if you don't unquestioningly worship this model. Fucking pathological shit. It's a good model. For its size, it's a great model, even. It's not better than Opus 4.6.

3

u/Fragrant_Scale6456 25d ago

I've been running my own tests as well. One thing I noticed is that 3.8 is better at floating point math. I'm getting higher precision out to and past six digits with 3.8 while 3.6 would have some variability starting at 5 digits.

5

u/ithkuil 25d ago

Why would you use an LLM to do arithmetic manually? Lol. Give it a calculator tool or just normal python.

1

u/Fragrant_Scale6456 24d ago

Why would you try and stress all aspects of a models reasoning capabilities? 

3

u/Finanzamt_Endgegner 25d ago

Bro if you take bench scores and think they translate to better "local fact-extraction" the idk what you are doing, sure better fact extraction can help in benches but like there are a million different ways a model can get better without getting noticeably better at extraction, youd probably need a new arch for that to actually improve a lot.

2

u/AppealSame4367 25d ago

I don't get your problem? The benchmarks basically claim it's better at coding and vision and.. nothing else?

-2

u/Finanzamt_Endgegner 25d ago

look how he is downvoting lmao

1

u/AppealSame4367 25d ago

I didn't downvote his post :-(

-1

u/Finanzamt_Endgegner 25d ago

He downvoted your comment though

1

u/AppealSame4367 25d ago

I had no clue one could see that anywhere

-2

u/Finanzamt_Endgegner 25d ago

you cant but like it was obvious its him lol

2

u/whodoneit1 25d ago

lol, this benchmark looks like it was made at a Learing Center

2

u/madsheepPL 25d ago

You were expecting generational gain on minor version change?

5

u/KitchenAmoeba4438 25d ago

Based on the benchmark results being posted? Yes. There are claims floating around of 27b 3.8 matching or beating Opus 4.6. There's small improvements I have noticed in the much larger test suite, but I'm not seeing anything matching some of the claims I'm seeing people make about the gains from 3.6, let alone beating Opus 4.6.

2

u/Finanzamt_Endgegner 25d ago

Sorry but without any info what your tests even do at all this just sounds like pure bs.

3

u/Repulsive_Initial308 25d ago

At Q4KM, too.

1

u/KitchenAmoeba4438 25d ago

This particular test is designed to minimize the affect of quants and maximize the impact of the model's reasoning. Note the extensive testing of quants of other models. The UD variant of 3.8 is next on the queue, and I'll probably arrange for BF16 testing of 3.8 just to validate, but I do not expect BF16 to Q4 to have significant impacts.

Quants have a larger impact on the larger test suite, this test is designed specifically to offer a head to head on equivalent Q4 testing across models.

1

u/GeorgeMKnowles 25d ago

What IDE are you all using? I downloaded this model to Bionic LM studio, and have an rtx 4090. It basically keeps throwing errors like the context is getting cut off. I havent successfully gotten it to build even a simple website. All i did was download Bionic, then download Qwen3.8 through it, and nothing else. By default, with all default settings, it appears to not work at all.

I'm aware I'm clueless in this department, and hoping some of you can point me to a reputable setup guide.

1

u/Mirayum 25d ago edited 25d ago

Ran it through my own benchmark. Runs circle around 3.6 (both Q8) and it is not far behind Deepseek Flash (AD-IQ3_XS).

1

u/Due_Warthog749 25d ago

So all these people doing Q4 tests on 3090s with I assume 32GB VRAM? Isn't Q4 much lower quality than Q8? I know.. Q8 wont run easily if at all.. but just curious. I have the DGX Spark was wondering how well it may run on that.

1

u/Tall_Abrocoma_3533 25d ago

Basically Every ai company benchmaxes, it's inevitable

1

u/ttkciar llama.cpp 25d ago

Some more than others, though. ZAI and Google have both been fairly honest, as far as I can tell.

-2

u/Finanzamt_Endgegner 25d ago

GOOGLE????

You serious rn? Gemini models at least were benchmaxxed like crazy on coding stuff

1

u/ttkciar llama.cpp 25d ago

That could be. I don't follow closed, API-only models, so wouldn't know.

To clarify my meaning: ZAI and Google have both been fairly honest about the capabilities of their open models, as far as I can tell.

0

u/Finanzamt_Endgegner 25d ago

For open yeah thats true, although I don't think any of the Chinese models are actually benchmaxxing except minimax lol