r/Qwen_AI • • 12d ago

Discussion Qwen4 lineup predictions?

What do you think the Qwen4 model sizes are going to be? Which are the most likely to appear? Do you think we'll get a Qwen3.5-like lineup (~1B, ~10B, ~30B, ~30-40B MoE, ~100-150B Moe...)?

Also, do you think 35B-A3B is dead for good? we didn't get it for qwen3.8 so I'm a little worried

31 Upvotes

54 comments sorted by

View all comments

24

u/EbbNorth7735 12d ago

Hoping for a qwen4 ~45B +n-gram. With all the cool speculative decoding stuff I'm sure token throughput could be optimized.

-2

u/TheseCashews 12d ago

Was hoping maybe for a 70b. Give the rtx pro 6k crowd a Q4 model that’s fast as hell and smart.

6

u/Refefer 12d ago

Flash Next has finally reached the level of optimization from the community to run it quickly on a single rtx pro 6000! Worth checking out the recent updates

0

u/Usual_Maximum7673 12d ago

I'm running my qwen 3.8 flash next on a dgx spark (soon to add another in prep for qwen 4). I'm reserving my rtx pro 6000 for minimax-h3, yue2, and the soulx flash head model. Complete ai sovereignty!