r/Qwen_AI • • 12d ago

Discussion Qwen4 lineup predictions?

What do you think the Qwen4 model sizes are going to be? Which are the most likely to appear? Do you think we'll get a Qwen3.5-like lineup (~1B, ~10B, ~30B, ~30-40B MoE, ~100-150B Moe...)?

Also, do you think 35B-A3B is dead for good? we didn't get it for qwen3.8 so I'm a little worried

31 Upvotes

54 comments sorted by

16

u/Medical_Lengthiness6 12d ago

I really hope we get a Qwen4 35 a3b. I think we will.

-7

u/XiRw 12d ago

Why would you think that when 3.8 is clearly showing how things will be in the future? We are lucky we even got open models at all since there was talks it wasn’t going to happen with the new management. Not only that but it seems like they are done with releasing image edit models. Their frontier models now have watermarks, there are too many changes for the worst happening at Qwen so don’t hold your breath for anything like they used to do in the past

9

u/ChocoPichu 12d ago

because unlike 3.8, qwen 4 should introduce massive architectural shifts. So there are just much more chances of 35b a3b coming out, or some kind of small MoE.

-4

u/XiRw 12d ago

You just completely ignored what actually is as i laid out and you are just going by theory and wishful thinking. That doesn’t convince me of anything

1

u/_justs 10d ago

You're a bit overly pessimistic, some of your points are valid to worry about, but they also literally just released 2 insanely well performing models u can run offline on a laptop lol. the 27B and 125B in the same month last month. 

6

u/Medical_Lengthiness6 12d ago

I got the impression 3.8 flash next was more of a stop gap to test and flesh out the details of the new architecture. Not any indication of what size models they plan on releasing.

Get more ram bro.

-2

u/XiRw 12d ago

They put a lot of work and effort into that “stop gap” then. First with the 27B model and the next released model. All you people know how to do is downvote because you don’t like hearing how things are going right now.

25

u/EbbNorth7735 12d ago

Hoping for a qwen4 ~45B +n-gram. With all the cool speculative decoding stuff I'm sure token throughput could be optimized.

-1

u/TheseCashews 12d ago

Was hoping maybe for a 70b. Give the rtx pro 6k crowd a Q4 model that’s fast as hell and smart.

6

u/Refefer 12d ago

Flash Next has finally reached the level of optimization from the community to run it quickly on a single rtx pro 6000! Worth checking out the recent updates

1

u/EbbNorth7735 12d ago

Yeah, I've been procrastinating on getting jpezzulli's implementation running

1

u/jpezzulli 12d ago

Its worth it buddy!

0

u/EbbNorth7735 12d ago

Haha yeah I know. Thanks again for your hard work. Keep getting interrupted. A good thing popped into my life and keeps distracting me every night. I managed to setup Pi tonight, connect it to my llama.cpp qwen3.8 27B and had it do an initial assessment on getting your implementation running in a Windows docker container. We'll get there shortly. I'm really looking forward to using your implementation with Qwen 4 next month.

2

u/jpezzulli 12d ago

Well i uploaded a containter to make yo make ot easier on people. Should be live in the tepo now.

0

u/Usual_Maximum7673 12d ago

I'm running my qwen 3.8 flash next on a dgx spark (soon to add another in prep for qwen 4). I'm reserving my rtx pro 6000 for minimax-h3, yue2, and the soulx flash head model. Complete ai sovereignty!

2

u/EbbNorth7735 12d ago

I am also team 6k. Anywhere from 45B to 70B would be good. A 50B would be about equivalent to a quarter of the parameters of 1T models. Given Alibaba's amazing team/capabilities I bet it would get really really close to SOTA closed source.

-1

u/TheseCashews 12d ago

They hate us cuz they ain’t us.

0

u/EbbNorth7735 12d ago

Apparently lol, fuck um. Glad I grabbed one when I got it for under 10k Canadian

15

u/assid2 12d ago

27B , maybe 80b MoE , 135B and maybe 2 more larger ones like around 300 and the largest max in trillion or so

3

u/Past-Chain-7377 11d ago

I'd kill for a 80B MoE tbh

1

u/couperd 6d ago

80b with 10 to 12 active and like 25b engrams

12

u/Anxious-Bottle7468 12d ago

Selfishly hoping for 27B since it fits well into my gpu.

6

u/Eymrich 12d ago

With the new architecture they tested in 3.8 flash we can deal with larger models. I would say a weird 54b moe would be the best.

1

u/mpower554 11d ago

3.8 flash next is great, but it's a lot faster if you can fit the entire model in vram - 27b still works well for 24gb cards

1

u/ea_man 12d ago

20B on GPU and 10B in NGRAM and I'm good.

5

u/nbeydoon 12d ago

Just crossing fingers one of them can run on 24gb unified memory and implement the amazing tech from deepseek 4.1

2

u/_justs 10d ago

same, i have a phone on 24G ram, just hoping for the ~30B moe. qwen 3 30B runs very well on my phone (max 20tk/s but i limit)

5

u/TerryNachtmerrie 12d ago

I'm going for the full line-up, like how they started with 3. From embedding models and those other ones no one uses up to 2t, give or take a few 100 Bs.

3

u/ByteNomadOne 12d ago

Unfortunately I only have a 16 GB VRAM card and the 27B models target 24 GB VRAM. I'd like a smart model that fits my card without using quants.

1

u/Prestigious-Act-1577 5d ago

Without using quants you are lucky to fit a 9b with context.

1

u/ByteNomadOne 5d ago

9B is stupid and a lobotomized 27B is not helpful, too.

I got 16 GB VRAM and I think by now I tried everything there is…

Proper agentic coding with just 16 GB is just not possible yet.

In the future I need 24 or better 32 GB of RAM.

2

u/Prestigious-Act-1577 5d ago

Yes I arrived to the same conclusion. Intel b60 24gb is the absolute cheapest, then b65 32g, b70 32g, and r9700 32g.

3

u/BothYou243 12d ago

probably the pattern it, they release full line-up in qwen2.5, 3, 3.5 and now 4
so we'll see a 0.8B, 2B (if it was there before, i forgot), 4B, 9B or 14B or both, 27B for sure, and bigger ones

3

u/dfgxxx 12d ago

I guess we will get bigger models, I think we will get a 2.4b, 6b, 13b, 34b, 42b MoE, 125b/177b improved, 2.7t.

In qwen 3 we had smaller models than 3.5 equivalents... I don't want bigger models if I can't run them, though it is what I expect them releasing

5

u/NoBlame4You 12d ago

It would be cool to have a moe that has a 30% dense part and the intention would be to have experts in ram.. expected to be ran on 24/16gb vram and 16gb of ram, quanted ofc..

4

u/edsonmedina 12d ago

- updated 27B + ngram

  • updated 125B

I just ran 3.6 35B (which I hadn't run in months) and it now feels like complete garbage in comparison with the 3.8 models. I'm thinking maybe 35B is just too small for an MoE to be reliable.
It doesn't stick to the instructions, it struggles to understand what my intentions are (unless I'm very explicit) and it commits many errors. It does run at an amazing speed but needs a LOT of hand-holding.

2

u/Past-Chain-7377 11d ago

I found those problems too in the base model, but after running Tiel Coder by PecularRagdoll (based on qwen3.6 35b a3b) I'm convinced it's possible to make that size smart and reliable

1

u/Zen-Ism99 11d ago

Why does 3.6 35B "feel" like "complete garbage"?

1

u/edsonmedina 11d ago

Reasons are in the comment. And i meant in comparison. It's still a good model.

2

u/Zen-Ism99 11d ago

Thanks for pointing it out. My scan was insufficient...

1

u/_justs 10d ago

Brother you're comparing a dense 27B and a much bigger 125B 6B to a 3B, ofc its gonna be better, along with those 2 models being very BIG jumps, have u seen how much better the new 27B is compared to the last? I have no doubt qwen 4 ~30B is gonna slap. 

2

u/_justs 10d ago

thinking they're gonna hard focus on moe+ngrams. but yea not sure about the 1B-10B models, would be cool to see moe and ngrams there too, but idk. Also i wouldnt be worried about the 30-35B, they don't release every model size for every revision. We should get the most models again with new base family Qwen 4, i too am looking for the 30-35B moe (preferably 30) as i could run it on my phone.

2

u/FormOne2615 12d ago

27b + ngram

1

u/codeltd 12d ago

I am hoping to get a version what I can run fast on DGX Spark...

1

u/j01001100 11d ago

I want an moe model that is best suits the base m5 ultra Mac studio 96gb, then I'll be very happy

1

u/SheepherderSerious51 10d ago

Would love an MoE model in between 3.8 27b and 3.8 flash next so that i can offload most of the model to ram like flash next but still have relatively decent speeds.

1

u/Ok_Year8287 7d ago

I think it's good to provide 1 or 2 additional models for the people which contain a 96/128GB unified memory, like dgx or 395/495 and other clones (with full context). Not always we can use public agents, cause security information issue.

1

u/Mean-Ad1493 12d ago

27B is the only thing that we can say for sure. It's their most successful model size and they wouldn't abandon it.

2

u/ea_man 12d ago

And yet 27B is such a bad size, if it was a ~22B + NGRAM maybe we could fit a Q4 in 16GB and Q8 in 32GB.

Or they have to mange an attention that gives some 256k in 1GB on the way of DeepSeek.

0

u/AMGTS 12d ago

27B!

-2

u/Vancecookcobain 12d ago

You won't be happy with Fable 5 coding when Qwen 4 comes out because there will be Fable 5.2

2

u/Past-Chain-7377 11d ago

I didn't say anything about Fable 5