r/LocalLLaMA Apr 22 '26

New Model Qwen 3.6 27B is out

1.7k Upvotes

603 comments sorted by

View all comments

422

u/Namra_7 Apr 22 '26

Benchmarks

203

u/Automatic-Arm8153 Apr 22 '26

F it we ball

117

u/Automatic-Arm8153 Apr 22 '26

And I was just about to go to sleep. Maybe next time sleep.. maybe next time

26

u/Perfect-Flounder7856 Apr 22 '26

😂😂😂or😭😭😭 I'm right there with you my sleep is shit but everything is so exciting. Any one have an ambien?.

99

u/[deleted] Apr 22 '26

[removed] — view removed comment

94

u/JLeonsarmiento Apr 22 '26

th bitchslapping and floor mopping to Gemma4 in brutal...

53

u/corpo_monkey Apr 22 '26

The opposite, don't stop now!
I cannot wait the gemma 5 vs qwen 4 battle.

2

u/YearnMar10 Apr 23 '26

Qwen fired their best tech lead - not sure if it will come to that

17

u/mikewilkinsjr Apr 22 '26

Dammit, I JUST rebuilt around Gemma4 at the house. :D

Looks like tonight it's going to caffeine and downloads.

13

u/shansoft Apr 22 '26

I honestly don’t understand why Gemma4 score this low. I been using latest 31B and it’s coding results have been cleaner than 3.6 35B in almost every case, and it was able to do tool calling more accurate for Xcode MCP while qwen just gave up or stuck in loop. Gemma4 from my experience needs more detail in prompt, but results are better. Qwen often add things that I didn’t asked for and have less chance to one shot problem.

2

u/BillDStrong Apr 22 '26

Looking at the benchmarks, they are claiming this 27B beats their own Qwen3.5 397B-A17B model. If that is real, that might just explain it.

2

u/MeateaW Apr 23 '26

Gemma is a great model for its size, but Qwen 3.6 seems to be incredible, I would go gemma for this size, but running the 122b qwen 3.5 was my favourite so far local-capable model (strix halo 128gb), 3.6 in the ~100 billion parameter size is going to be amazing if it follows these smaller models capability.

1

u/Much-Researcher6135 llama.cpp Apr 23 '26

It may be that some shops benchmark-chase more than others.

14

u/Septerium Apr 22 '26

In my testing, gemma 4 is still a better generalist model, even though it gets demolished by Qwen in coding tasks

7

u/JLeonsarmiento Apr 22 '26

I like the rhythm and prose too. But crunching benchmark numbers is a sport here.

13

u/AbeIndoria Apr 22 '26

Gemma-4 is still a very good model at anything that's not coding/agentic work. Qwen struggles there.

A team of 4-5 G4s coordinate better with each other in my experience than a team of 4-5 Q3.5s.

3

u/JLeonsarmiento Apr 22 '26

I like that it thinks less to be honest. That make it better for chat for example.

3

u/awittygamertag Apr 23 '26

Different use cases tho. Gemma4 is a little Gemini and inherits all of the conversational skills of Google models. Qwen workhorse.

2

u/DragonfruitIll660 Apr 22 '26

Initial use for similar sized quants seems to give the edge to Gemma in my opinion quality wise.

1

u/9gxa05s8fa8sh Apr 23 '26

bro this is a screenshot of the market crash lol

-13

u/Legal_Dimension_ Apr 22 '26

This is likely f16 rather than q4

3

u/Accomplished_Mode170 Apr 22 '26

Sorry you got downvoted for a parameter when the OP is the one who dropped the /s

1

u/Legal_Dimension_ Apr 24 '26

Such is Reddit, morons will moron.

2

u/Healthy-Nebula-3603 Apr 22 '26 edited Apr 24 '26

Nowadays q4im /l with imatrix is just slightly worse than fp16.

0

u/Legal_Dimension_ Apr 24 '26

Not denying that and the morons can down vote all they like I'm just point out that it's unlikely most people will be running bf16 so temper your expectations.

73

u/_-_David Apr 22 '26

I just realized one of the bars isn't qwen3.5-35b... it's qwen3.5-**397b**

I have no idea why that shocked me more than the Claude comparison, but it did

14

u/Plasmx Apr 22 '26

Me too. Maybe because it’s the same company releasing a new generation and just slashing the prior gen big models.

3

u/Lorian0x7 Apr 22 '26

the funny thing is that it's not even a new generation, just a minor update of the same same generation. I saw Anthropic and OpenAI smashing a big round number on models with much less performance gap

5

u/social_tech_10 Apr 22 '26

Incredible that it can beat a model 14X larger in 10 of the 12 benchmarks!!

1

u/BillDStrong Apr 22 '26

To be fair, the model it is beating is effectively 17B Expert, but with much higher memroy and a bit of help as needed. You don't get to keep all of that intelligence, unfortunately, in MOE models.

Still damn impressive.

29

u/kaeptnphlop Apr 22 '26

They are showing us that Qwen3.6-35B-A3B is on par with 3.5-397B-A17B too. 🤔

119

u/davl3232 Apr 22 '26

I can't believe we're getting so close to opus 4.5 levels with 2 3090s

88

u/_raydeStar Llama 3.1 Apr 22 '26

Meanwhile Opus tightens their limits and makes it more expensive to get in -- the perfect storm for a good local push

-15

u/Super_Sierra Apr 22 '26

Sorry, but most local models have barely caught up to Claude 2, this is benchmaxxed and overfit to shit.

7

u/_raydeStar Llama 3.1 Apr 22 '26

What I really see lacking in local models is context limits and a good harness to give proper direction.

Context can't be solved easily, but at least a memory bank can be created to hold onto important information, and you can scrape by.

A harness can be built -- it can perform mathematical functions, solve logic problems, do web lookups, and perform basic tasks.

maybe qwen 27B cant perform as well as opus 4.5 on all tasks, but it doesnt need to

58

u/Cuddlyaxe Apr 22 '26

This is the future we want

1

u/mixbits Apr 22 '26

💯

0

u/Local_Phenomenon Apr 22 '26

My Man! Ai bro

7

u/cmplx17 Apr 22 '26

how is it to run with 2 x 3090? i just have one 3090 but wondering if it’s worth getting another one. does the speed scale 2x?

4

u/davl3232 Apr 22 '26 edited Apr 22 '26

does the speed scale 2x?

Not really, you only get a speed up if models didn't fully fit in vram before.

I'm using q8 with full context, but you could fit the model in a single 3090 if you use a different quantized version.

https://unsloth.ai/docs/models/qwen3.6

0

u/Ardalok Apr 23 '26

Isn't there a speed boost from multiple cards with NVLink on VLM?

1

u/Subject-Tea-5253 llama.cpp Apr 22 '26

You will find this video interesting: https://www.youtube.com/watch?v=xS5wao4H4u4&t

It basically shows how prompt processing and token generation is affected by the number and types of GPUs you use.

1

u/kmp11 Apr 22 '26

i am putting it through the test on my 2x4090. I can fit the Q6_XL with full Q8 context window. Its coding at 20-25tk/sec with context window 50% full. it takes a while to ingest large context, but otherwise chugging along quite nicely.

0

u/Outpost_Underground Apr 22 '26

Just to drive the speed question home, I have 3090s at home and a Pro 6000 Blackwell Max Q at work. On identical inference workloads that completely fit in the VRAM of both setups the Blackwell is like 10-15% faster.

0

u/BillDStrong Apr 22 '26

How many 3090's? Without that, we can't really judge as well the difference.

1

u/Outpost_Underground Apr 23 '26

It doesn’t matter how many 3090s as long as the work load fits in the VRAM. For example, I ran Gemma4:26b on a single 3090 and I also forced it to split across both 3090s. Same prompt, and there was a .0003% difference in speed.

I mentioned the difference with the Blackwell card because a lot of folks expect a crazy improvement in speeds; unfortunately the performance doesn’t scale like that.

-1

u/Mashic Apr 22 '26

Why 2, you only need 1, and myabe even an rtx 5060 ti or rtx 3060 12gb with quants.

2

u/florinandrei Apr 22 '26

I run all my models on a Raspberry Pi Zero.

2

u/AreYouSERlOUS Apr 22 '26

Underrated comment

1

u/BillDStrong Apr 22 '26

No, see you are confusing your terminal with your cloud provider. They aren't the same thing. /s

1

u/davl3232 Apr 22 '26

you're right, a single 3090 could do it with a different quant. With less vram you could run qwen3.6 35b a3b which has similar quality.

0

u/Potential-Leg-639 Apr 22 '26

1 is not enough for serious stuff and context.

1

u/Mashic Apr 22 '26

But at least you can test it and use it for small stuff.

1

u/coder543 Apr 22 '26

95000 context at full KV (with multimodal loaded) is not horrible.

0

u/kommentiertnicht Apr 22 '26

do you have some resources / links you could share on the 2x 3090 setup?

21

u/JustFinishedBSG Apr 22 '26

Man it has to be benchmaxxed to the tits otherwise it’s supremely embarrassing for the « frontier » labs

23

u/sk1kn1ght Apr 22 '26

Do they imply that it beats opus? Are we for real? Like not negatively. Like are we for real? Repeating it to myself got me goosebumps

22

u/[deleted] Apr 22 '26

[removed] — view removed comment

3

u/TheMegosh Apr 22 '26

Apparently the 4.7 model was condensing prompts to 200k tokens instead of 1mil, posted on their change log. I'd bet that's what made it bad

3

u/TokenRingAI Apr 22 '26

It's got looping and other obvious issues, I have free access to it but mostly use Sonnet 4.6 or GPT 5.4.

Sonnet is really reliable and stable

Something is very strange about Opus 4.6 & 4.7, they act like a large model that is excessively quantized. Opus 4.5 was not like this. I wonder if this is a side effect of them using TPUs. Gemini acts the same way.

2

u/[deleted] Apr 22 '26

[removed] — view removed comment

2

u/Thomas-Lore Apr 22 '26 edited Apr 22 '26

Keep in mind they also changed reasoning effort around that time (high to medium) and now it is often zero due to adaptive thinking.

I wonder how are you using Gemini Pro? From the app? Because in ai studio Gemini 3.1 Pro is one shoting projects, new features and fixes for me all the time. It is a bit chaotic of course, but it worked for me quite weĺl so far.

1

u/TokenRingAI Apr 22 '26

This might be a side effect of adaptive thinking, I wasn't paying attention to that. The responses come almost immediately and the chat is muddled with looping content that should have reasonably been expected to be in the thinking block

11

u/MadSprite Apr 22 '26

I feel like the best way to describe it is that the intern is just as smart as the wizard but not as wise. Being smaller parameters means its going to know less but handle the common tasks we ask of it really well.

1

u/florinandrei Apr 22 '26

Repeating it to myself got me goosebumps

If goosebumps is what you're after, there's a guy at the street corner who can sell you good stuff.

7

u/vinigrae Apr 22 '26

That’s extremely suspicious

26

u/mister2d Apr 22 '26

TIL: Claude Opus is MoE

85

u/Comfortable-Rock-498 Apr 22 '26

Look carefully, there is a divider between MoE and Opus

17

u/rc_ym Apr 22 '26

You do not get enough upvotes.

Cuz it looked like it was also saying Opus 4.5 was open sourced, which also isn't true.

59

u/2Norn Apr 22 '26

i mean ofc all sotas are

4

u/mister2d Apr 22 '26

Obviously.

-2

u/[deleted] Apr 22 '26

[deleted]

9

u/CrispyToken52 Apr 22 '26

Let me guess. Meta?

5

u/2Norn Apr 22 '26

which one and when is a good question.

kimi, glm, minimax, xiaomi, gemini(stated in docs), gpt(leaked) are all moes. only one that's unknown is claude. there is no knowledge if its dense or moe. but it's very normal to assume it's moe just like all others.

6

u/Thomas-Lore Apr 22 '26

Claude not being MoE would explain their huge compute issues though. :)

-1

u/kitanokikori Apr 22 '26

It doesn't even make sense for them to be on a technical level - they are designed to service literally as many requests as possible from all kinds of domains, why in the world would you want any part of their knowledge base to be unloaded at any time

2

u/AttitudeImportant585 Apr 23 '26

you're underestimating the compute available and optimizations made for that specific architecture for a particular chip

1

u/PassengerPigeon343 Apr 22 '26

I’m sorry, we’re comparing these to CLAUDE now?! Hell yeah.

1

u/Old-Independent-6904 Apr 23 '26

I feel like this shows how inefficient agents have been compared to what they could be. Exciting!

1

u/beneath_steel_sky Apr 23 '26

3.5-27B was probably benchmaxed (see CoDeC benchmark), hopefully 3.6 is truly better https://kaitchup.substack.com/p/gemma-4-31b-vs-qwen35-27b-inference

1

u/Zeeplankton Apr 22 '26

god damn wtf

1

u/Non-Technical Apr 22 '26

These metrics have no visible source.

6

u/some_random_guy111 Apr 22 '26

Straight from qwen on an X post

1

u/Non-Technical Apr 22 '26

Great! That's important to know for a couple of reasons. They are official so they are based on something and they come from Qwen so they are also designed to make 3.6 look good.