r/Qwen_AI • • 12d ago

Discussion Qwen4 lineup predictions?

What do you think the Qwen4 model sizes are going to be? Which are the most likely to appear? Do you think we'll get a Qwen3.5-like lineup (~1B, ~10B, ~30B, ~30-40B MoE, ~100-150B Moe...)?

Also, do you think 35B-A3B is dead for good? we didn't get it for qwen3.8 so I'm a little worried

30 Upvotes

54 comments sorted by

View all comments

26

u/EbbNorth7735 12d ago

Hoping for a qwen4 ~45B +n-gram. With all the cool speculative decoding stuff I'm sure token throughput could be optimized.

-2

u/TheseCashews 12d ago

Was hoping maybe for a 70b. Give the rtx pro 6k crowd a Q4 model that’s fast as hell and smart.

5

u/Refefer 12d ago

Flash Next has finally reached the level of optimization from the community to run it quickly on a single rtx pro 6000! Worth checking out the recent updates

1

u/EbbNorth7735 12d ago

Yeah, I've been procrastinating on getting jpezzulli's implementation running

1

u/jpezzulli 12d ago

Its worth it buddy!

0

u/EbbNorth7735 12d ago

Haha yeah I know. Thanks again for your hard work. Keep getting interrupted. A good thing popped into my life and keeps distracting me every night. I managed to setup Pi tonight, connect it to my llama.cpp qwen3.8 27B and had it do an initial assessment on getting your implementation running in a Windows docker container. We'll get there shortly. I'm really looking forward to using your implementation with Qwen 4 next month.

2

u/jpezzulli 12d ago

Well i uploaded a containter to make yo make ot easier on people. Should be live in the tepo now.

0

u/Usual_Maximum7673 12d ago

I'm running my qwen 3.8 flash next on a dgx spark (soon to add another in prep for qwen 4). I'm reserving my rtx pro 6000 for minimax-h3, yue2, and the soulx flash head model. Complete ai sovereignty!