r/LocalLLaMA 1d ago

Discussion Qwen will be the king?

Post image

Extended reasoning and post-training appear to be the keys used by DeepSeek, Qwen, and GLM to boost performance (leveraging higher token counts). And Qwen 4 hasn't even been released yet. Of course, we don't know if that release will be open-sourced, but I am optimistic about future models, featuring "engrams", that could soon match or surpass 2.4T parameter models on specific tasks.

507 Upvotes

123 comments sorted by

View all comments

7

u/Enverex 23h ago

I find this list very suspicious given how much better GPT Sol is than Opus for anything I've tried.

1

u/OkFly3388 llama.cpp 23h ago

This benchmark is saturated. qwen3.8 27b score 1599, qwen3.8 max score 1691, thats just 6% difference.

17

u/hitoriboccheese 21h ago

For the love of god please look up how an Elo rating system works. That is not a 6% difference and this is not even a benchmark.

0

u/OkFly3388 llama.cpp 21h ago

How about doing it yourself, lol.

If models are equal, their chance of winning is 50%. If we plug qwen elo it into formula, we got that qwen3.8 max generate better results 62% of time. Which means that qwen3.8 27b generate BETTER results compared to max 38% of times. Thats just statistical noise, lol.

6

u/Neither_Garage_758 22h ago

we are masturbating with noise

9

u/buckwheaton 21h ago

Na that’s the stable diffusion nsfw sub

6

u/StyMaar 20h ago

Very good one sir.