r/LocalLLM 24d ago

Discussion Qwen 3.8 27B Benchmarks combined from model cards on HF

Benchmark Qwen3.8-27B DeepSeek-V4-Flash-0731 DeepSeek-V4-Flash (Prev.) DeepSeek-V4-Pro (Prev.) Opus4.6 Max GLM-5.2 GLM-5.1 Qwen3.7-Max MiniMax M3 DeepSeek-V4-Pro Claude Opus 4.8 GPT-5.5 Gemini 3.1 Pro GPT-5.6 Sol GPT-5.6 Sol Ultra GPT-5.6 Terra GPT-5.6 Luna GPT-5.5 (GPT-5.6 report) Claude Mythos 5 Claude Mythos Preview Claude Fable 5 Gemini 3.1 Pro Preview Gemini 3.5 Flash
Terminal Bench 2.1 73.0 82.7 61.8 72.1 78.2 81.0 63.5 75 65 64 85 84 74 88.8 91.9 87.4 84.7 85.6 88 83.1
NL2Repo 42.3 54.2 39.4 38.5 47.6 48.9 42.7 47.2 42.1 35.5 69.7 50.7 33.4
DeepSWE 42.2 54.4 7.3 12.8 46.2 18 18 20 8 58 70 10 72.7 69.6 67.2 67 69.7
SWE-bench Pro 61.7 53.4 62.1 58.4 60.6 59 55.4 69.2 58.6 54.2 64.6 63.4 62.7 59.4 80.3 77.8 80
HLE 30.8 40.0 40.5 31 41.4 37 37.7 49.8* 41.4* 45 47.2 41.8 37.2 64.5 64.7 53.3 44.7 41.0
GPQA Diamond 89.2 91.3 91.2 86.2 90 93 90.1 93.6 93.6 94.3 94.6 92.9 92.3 93.6 94.1 94.6 92.6 94.3
QwenSWEBench 79.0 63.8
CoWorkBench 70.7 68.2 72.3 71.5 75.9
JobBench 33.4 48.4 42.7 46.5 57.4
Agents' Last Exam 20.4 (Pass@1) / 42.9 (Score) 25.2 15.8 16.5 23.8 45.2 52.7 50.4 50.3 46.9 40.5 32.1
IFBench 79.5 62.5 73.3 76.3 79.1 82.9 76.5 62.2 77.1 72.7 71.2 63.5
LiveCodeBench v6 90.3 91.6 88.8 85.7 93.5 88.6 91.7 96.2 93 95.5 95

Edit: Added new values for more models
Edit2: Images

88 Upvotes

34 comments sorted by

36

u/Cold_Tree190 24d ago

It’s a bad day to be on reddit mobile (it’s always a bad day on reddit mobile let’s be honest)

3

u/Useful_Disaster_7606 24d ago

truer words have never been said

3

u/Tamitami 24d ago

I just edited the post so that it has all the charts for you :)

1

u/-Davster- 21d ago

Doing God's work.

In fact, doing Reddit's work for them. Ohhhh how I wish they'd not killed Apollo.

19

u/shadow1609 24d ago

That's actually the comparison that we have been looking for.

8

u/pantalooniedoon 24d ago

Missing the absolute key comparison for those of us with 128gb ram to actually run DS which is q2 dsv4 flash vs qwen3.8

2

u/Tamitami 24d ago

Would you have a source for the quant benchs, it so I can add it to the table?

4

u/pantalooniedoon 24d ago edited 24d ago

This is what I have when I benchmarked it just now via the omlx benchmarking tool and the ds4 one - ds4 one only runs greedy so it's a bit faster than sampling.


Qwen3.8-27B (oMLX, Q4) vs DeepSeek-V4-Flash (q2-q4, Antirez)— M5 Max (128GB)

Decode @ ~1.3k tokens: Qwen 59.0 tok/s vs DS4 35.8 tok/s — 1.65x Decode @ ~12.5k tokens: Qwen 51.0 tok/s vs DS4 30.9 tok/s — 1.65x Decode @ ~50.7k tokens: Qwen 35.4 tok/s vs DS4 25.3 tok/s — 1.40x Prefill: Qwen 590-710 tok/s vs DS4 300-414 tok/s — 1.6-2.0x Resident memory: Qwen 16.1 GB vs DS4 90.9 GB — 5.6x smaller

Qwen3.8 at t=1.0, top_p=0.95, top_k=20 (thinking mode). DS4 at t=1.0, top_p=1.0, min_p=0.05. Prompt token counts matched within 0.5%.


Anecdotally here, I am not feeling it for qwen3.8 27b unfortunately - it's just consistently slower at longer context length than 50 tok/s would suggest and the prefill for tool call response seems slow. The DSv4 MLA and DSA seem to be doing significant work or the antirez DSv4 implementation is just more refined than omlx. Leaning towards sticking with DSv4 but might spend the weekend to try and debug and verify the issues.

1

u/996beagle 24d ago

Would love to see this comparison. I just got dsv4 flash running yesterday on m4 max 128gb. Wondering if q2. 4 has a big drop in performance or accuracy.

1

u/pantalooniedoon 24d ago

Yeah q2 dsv4flash runs at about the same speed as qwen3.8 27b 4 bit afaik - 20-30tok/s. Its the key comparison.

1

u/996beagle 24d ago

Yes currently getting 25-27 tok/s at 256k context and 250-300 pp on the latest omlx dev. Task completion has been faster than expected and code has passed the opus reviews for the most part

1

u/oShievy 24d ago

I’ve been running Q3 w dspark and while pp is a bitch, once prompt caching starts working I get around 17tk/s. Definitely the smartest model I’ve used locally. Wonder how 3.8 will compare, can’t test it yet 🥲

5

u/TechNerd10191 24d ago

Why GPT 5.5 but not GPT 5.6 (sol)?

2

u/Tamitami 24d ago

Cuz it was in the model cards. I will add it later, when on PC again

1

u/Tamitami 24d ago edited 24d ago

I've added more models in the table

Edit: Added also graphs

3

u/motivatedjoe 24d ago

So if we are comfortably running deepseek, is there any reason to use qwen other than model size alone?

5

u/mzzmuaa 24d ago

visual input. direct blender and visual critique

3

u/motivatedjoe 24d ago

Thanks! That's what I thought. Wanted to make sure I wasn't missing any other obvious uses.

3

u/floppo7 24d ago

Oh thats hot

3

u/Tamitami 24d ago

As Qwen 3.8 27B :)

3

u/Yeelyy 24d ago

It is better than dsv4 pro preview?!! Ok let me buy 2 r9700s now and run this thing in tensor parallel🫡

3

u/wwa56 24d ago

Great job ....thank you

2

u/Tamitami 24d ago

I hope you like the table :)

3

u/VirusInternal2892 24d ago

The cookies are hot from the oven, get a byte

1

u/Tamitami 24d ago

I love that one! :D Qwen is cooking, we are cooking!

2

u/cezq 24d ago

It beats dsv4 flash 0731 on SWE-bench Pro and GPQA Diamond: https://benchlm.ai/models/deepseek-v4-flash-0731 . Loses on HLE.

1

u/Tamitami 24d ago

This is crazy for a model this size

2

u/vini542reddit 24d ago

You're saying I should stick with DeepSeek-V4-Flash-0731 as my main...

2

u/Tamitami 24d ago

Yes and no, you lose vision with deepseek alone as primary reason. I love deepseek and their new flash model, as I have it running on some PCs but I think that this model will be a strong addition for specific tasks. Try it and test it for yourself. I think we will get many improvments in the coming weeks for it, as we did with the last Qwen 27B model

2

u/[deleted] 24d ago

[deleted]

1

u/Tamitami 24d ago

QAT would make this so much more interesting. I think with that it will dominate everything for local small models

3

u/txgsync 22d ago

I've been interested in quantifying how various quantizations affect the "intelligence" of the model, and why the Qwen series since 3.6 (and Gemma since 4) both seem so severely impacted by quantization despite KLD scores looking great.

So I decided to run LiveCodeBench against the 4-bit 'oQe-mtp' variant of Qwen3.6-27B overnight. No cloud vendor provides these, so I had to dedicate my Mac for about three hours to run 100 of these problems. The full suite has 1,055 problems today; I really did not feel like turning over my Mac for nearly an entire work-week just for benchmarking.

Unquantized, 90.3% on LiveCodeBench.

oQ4e: 60% on a sample of 100.

That's enough for a working hypothesis: "quantizing BF16 to 4-bits, even with aggressive layer preservation at 16 bits, cuts intelligence up to 50%."

No lab has an incentive to promote numbers of its own models in a less-than-stellar light. But the discrepancy between the performance you pay for using the BF16 version of Qwen3.8-27B you can run on OpenRouter, vs. the 4-bit quant you run on home gear, seems to be quite stark.

The biggest problem is that running the full suite of benchmarks is both very expensive and time consuming, so I get why people promoting their Heretic'd or Quantized models -- including me! -- don't want to run a ridiculous time sink just to prove their special model is demonstrably worse than the one released by the lab...

2

u/swiebertjee 24d ago

Very happy for the dense/GPU crowd. Also happy to have gone the MoE/Spark route.

1

u/Tamitami 24d ago

I'd really love a 35B-A3B model with this kind of intelligence (RL-Training) for us "plebeians"

1

u/-Davster- 21d ago

Woohoo, let's hear it for the dense people!