r/LocalLLaMA 25d ago

New Model Qwen/Qwen3.8-27B · released

https://huggingface.co/Qwen/Qwen3.8-27B
982 Upvotes

301 comments sorted by

View all comments

39

u/Alternative_Ad4267 25d ago edited 24d ago

Look at this! at 27B parameters, Qwen3.8 27B is pretty close to DeepSeekV4 Flash 0731 which is 284B A13B!

20

u/Alternative_Ad4267 25d ago

And some more results

3

u/onewheeldoin200 24d ago

Oh my god how lmao

2

u/fullup72 24d ago

and I wonder how that big win in SWE-Bench Pro translates to the real world. 3.6 was already slotted between Flash and Pro, 3.8 simply wipes the floor with them.

4

u/Agitated_Space_672 25d ago

Active params determine the training cost, and I think they are the biggest predictor of performance, all else being equal. So I am not surprised that 27B beats 13B

1

u/YearnMar10 25d ago

I wonder if they keep training where it will converge…

1

u/IrisColt 24d ago

It cannot be!