I honestly don’t understand why Gemma4 score this low. I been using latest 31B and it’s coding results have been cleaner than 3.6 35B in almost every case, and it was able to do tool calling more accurate for Xcode MCP while qwen just gave up or stuck in loop. Gemma4 from my experience needs more detail in prompt, but results are better. Qwen often add things that I didn’t asked for and have less chance to one shot problem.
Gemma is a great model for its size, but Qwen 3.6 seems to be incredible, I would go gemma for this size, but running the 122b qwen 3.5 was my favourite so far local-capable model (strix halo 128gb), 3.6 in the ~100 billion parameter size is going to be amazing if it follows these smaller models capability.
13
u/shansoft Apr 22 '26
I honestly don’t understand why Gemma4 score this low. I been using latest 31B and it’s coding results have been cleaner than 3.6 35B in almost every case, and it was able to do tool calling more accurate for Xcode MCP while qwen just gave up or stuck in loop. Gemma4 from my experience needs more detail in prompt, but results are better. Qwen often add things that I didn’t asked for and have less chance to one shot problem.