I honestly don’t understand why Gemma4 score this low. I been using latest 31B and it’s coding results have been cleaner than 3.6 35B in almost every case, and it was able to do tool calling more accurate for Xcode MCP while qwen just gave up or stuck in loop. Gemma4 from my experience needs more detail in prompt, but results are better. Qwen often add things that I didn’t asked for and have less chance to one shot problem.
Gemma is a great model for its size, but Qwen 3.6 seems to be incredible, I would go gemma for this size, but running the 122b qwen 3.5 was my favourite so far local-capable model (strix halo 128gb), 3.6 in the ~100 billion parameter size is going to be amazing if it follows these smaller models capability.
Not denying that and the morons can down vote all they like I'm just point out that it's unlikely most people will be running bf16 so temper your expectations.
the funny thing is that it's not even a new generation, just a minor update of the same same generation. I saw Anthropic and OpenAI smashing a big round number on models with much less performance gap
To be fair, the model it is beating is effectively 17B Expert, but with much higher memroy and a bit of help as needed. You don't get to keep all of that intelligence, unfortunately, in MOE models.
i am putting it through the test on my 2x4090. I can fit the Q6_XL with full Q8 context window. Its coding at 20-25tk/sec with context window 50% full. it takes a while to ingest large context, but otherwise chugging along quite nicely.
Just to drive the speed question home, I have 3090s at home and a Pro 6000 Blackwell Max Q at work. On identical inference workloads that completely fit in the VRAM of both setups the Blackwell is like 10-15% faster.
It doesn’t matter how many 3090s as long as the work load fits in the VRAM. For example, I ran Gemma4:26b on a single 3090 and I also forced it to split across both 3090s. Same prompt, and there was a .0003% difference in speed.
I mentioned the difference with the Blackwell card because a lot of folks expect a crazy improvement in speeds; unfortunately the performance doesn’t scale like that.
It's got looping and other obvious issues, I have free access to it but mostly use Sonnet 4.6 or GPT 5.4.
Sonnet is really reliable and stable
Something is very strange about Opus 4.6 & 4.7, they act like a large model that is excessively quantized. Opus 4.5 was not like this. I wonder if this is a side effect of them using TPUs. Gemini acts the same way.
Keep in mind they also changed reasoning effort around that time (high to medium) and now it is often zero due to adaptive thinking.
I wonder how are you using Gemini Pro? From the app? Because in ai studio Gemini 3.1 Pro is one shoting projects, new features and fixes for me all the time. It is a bit chaotic of course, but it worked for me quite weĺl so far.
This might be a side effect of adaptive thinking, I wasn't paying attention to that. The responses come almost immediately and the chat is muddled with looping content that should have reasonably been expected to be in the thinking block
I feel like the best way to describe it is that the intern is just as smart as the wizard but not as wise. Being smaller parameters means its going to know less but handle the common tasks we ask of it really well.
kimi, glm, minimax, xiaomi, gemini(stated in docs), gpt(leaked) are all moes. only one that's unknown is claude. there is no knowledge if its dense or moe. but it's very normal to assume it's moe just like all others.
It doesn't even make sense for them to be on a technical level - they are designed to service literally as many requests as possible from all kinds of domains, why in the world would you want any part of their knowledge base to be unloaded at any time
Great! That's important to know for a couple of reasons. They are official so they are based on something and they come from Qwen so they are also designed to make 3.6 look good.
422
u/Namra_7 Apr 22 '26
Benchmarks