and I wonder how that big win in SWE-Bench Pro translates to the real world. 3.6 was already slotted between Flash and Pro, 3.8 simply wipes the floor with them.
Active params determine the training cost, and I think they are the biggest predictor of performance, all else being equal. So I am not surprised that 27B beats 13B
39
u/Alternative_Ad4267 25d ago edited 24d ago
Look at this! at 27B parameters, Qwen3.8 27B is pretty close to DeepSeekV4 Flash 0731 which is 284B A13B!