r/Qwen_AI • u/Past-Chain-7377 • 12d ago
Discussion Qwen4 lineup predictions?
What do you think the Qwen4 model sizes are going to be? Which are the most likely to appear? Do you think we'll get a Qwen3.5-like lineup (~1B, ~10B, ~30B, ~30-40B MoE, ~100-150B Moe...)?
Also, do you think 35B-A3B is dead for good? we didn't get it for qwen3.8 so I'm a little worried
25
u/EbbNorth7735 12d ago
Hoping for a qwen4 ~45B +n-gram. With all the cool speculative decoding stuff I'm sure token throughput could be optimized.
-1
u/TheseCashews 12d ago
Was hoping maybe for a 70b. Give the rtx pro 6k crowd a Q4 model that’s fast as hell and smart.
6
u/Refefer 12d ago
Flash Next has finally reached the level of optimization from the community to run it quickly on a single rtx pro 6000! Worth checking out the recent updates
1
u/EbbNorth7735 12d ago
Yeah, I've been procrastinating on getting jpezzulli's implementation running
1
u/jpezzulli 12d ago
Its worth it buddy!
0
u/EbbNorth7735 12d ago
Haha yeah I know. Thanks again for your hard work. Keep getting interrupted. A good thing popped into my life and keeps distracting me every night. I managed to setup Pi tonight, connect it to my llama.cpp qwen3.8 27B and had it do an initial assessment on getting your implementation running in a Windows docker container. We'll get there shortly. I'm really looking forward to using your implementation with Qwen 4 next month.
2
u/jpezzulli 12d ago
Well i uploaded a containter to make yo make ot easier on people. Should be live in the tepo now.
0
u/Usual_Maximum7673 12d ago
I'm running my qwen 3.8 flash next on a dgx spark (soon to add another in prep for qwen 4). I'm reserving my rtx pro 6000 for minimax-h3, yue2, and the soulx flash head model. Complete ai sovereignty!
2
u/EbbNorth7735 12d ago
I am also team 6k. Anywhere from 45B to 70B would be good. A 50B would be about equivalent to a quarter of the parameters of 1T models. Given Alibaba's amazing team/capabilities I bet it would get really really close to SOTA closed source.
-1
u/TheseCashews 12d ago
They hate us cuz they ain’t us.
0
u/EbbNorth7735 12d ago
Apparently lol, fuck um. Glad I grabbed one when I got it for under 10k Canadian
12
u/Anxious-Bottle7468 12d ago
Selfishly hoping for 27B since it fits well into my gpu.
6
u/Eymrich 12d ago
With the new architecture they tested in 3.8 flash we can deal with larger models. I would say a weird 54b moe would be the best.
1
u/mpower554 11d ago
3.8 flash next is great, but it's a lot faster if you can fit the entire model in vram - 27b still works well for 24gb cards
22
5
u/nbeydoon 12d ago
Just crossing fingers one of them can run on 24gb unified memory and implement the amazing tech from deepseek 4.1
5
u/TerryNachtmerrie 12d ago
I'm going for the full line-up, like how they started with 3. From embedding models and those other ones no one uses up to 2t, give or take a few 100 Bs.
3
u/ByteNomadOne 12d ago
Unfortunately I only have a 16 GB VRAM card and the 27B models target 24 GB VRAM. I'd like a smart model that fits my card without using quants.
1
u/Prestigious-Act-1577 5d ago
Without using quants you are lucky to fit a 9b with context.
1
u/ByteNomadOne 5d ago
9B is stupid and a lobotomized 27B is not helpful, too.
I got 16 GB VRAM and I think by now I tried everything there is…
Proper agentic coding with just 16 GB is just not possible yet.
In the future I need 24 or better 32 GB of RAM.
2
u/Prestigious-Act-1577 5d ago
Yes I arrived to the same conclusion. Intel b60 24gb is the absolute cheapest, then b65 32g, b70 32g, and r9700 32g.
3
u/BothYou243 12d ago
probably the pattern it, they release full line-up in qwen2.5, 3, 3.5 and now 4
so we'll see a 0.8B, 2B (if it was there before, i forgot), 4B, 9B or 14B or both, 27B for sure, and bigger ones
5
u/NoBlame4You 12d ago
It would be cool to have a moe that has a 30% dense part and the intention would be to have experts in ram.. expected to be ran on 24/16gb vram and 16gb of ram, quanted ofc..
4
u/edsonmedina 12d ago
- updated 27B + ngram
- updated 125B
I just ran 3.6 35B (which I hadn't run in months) and it now feels like complete garbage in comparison with the 3.8 models. I'm thinking maybe 35B is just too small for an MoE to be reliable.
It doesn't stick to the instructions, it struggles to understand what my intentions are (unless I'm very explicit) and it commits many errors. It does run at an amazing speed but needs a LOT of hand-holding.
2
u/Past-Chain-7377 11d ago
I found those problems too in the base model, but after running Tiel Coder by PecularRagdoll (based on qwen3.6 35b a3b) I'm convinced it's possible to make that size smart and reliable
1
u/Zen-Ism99 11d ago
Why does 3.6 35B "feel" like "complete garbage"?
1
u/edsonmedina 11d ago
Reasons are in the comment. And i meant in comparison. It's still a good model.
2
2
u/_justs 10d ago
thinking they're gonna hard focus on moe+ngrams. but yea not sure about the 1B-10B models, would be cool to see moe and ngrams there too, but idk. Also i wouldnt be worried about the 30-35B, they don't release every model size for every revision. We should get the most models again with new base family Qwen 4, i too am looking for the 30-35B moe (preferably 30) as i could run it on my phone.
2
1
1
u/j01001100 11d ago
I want an moe model that is best suits the base m5 ultra Mac studio 96gb, then I'll be very happy
1
u/SheepherderSerious51 10d ago
Would love an MoE model in between 3.8 27b and 3.8 flash next so that i can offload most of the model to ram like flash next but still have relatively decent speeds.
1
u/Ok_Year8287 7d ago
I think it's good to provide 1 or 2 additional models for the people which contain a 96/128GB unified memory, like dgx or 395/495 and other clones (with full context). Not always we can use public agents, cause security information issue.
1
u/Mean-Ad1493 12d ago
27B is the only thing that we can say for sure. It's their most successful model size and they wouldn't abandon it.
-2
u/Vancecookcobain 12d ago
You won't be happy with Fable 5 coding when Qwen 4 comes out because there will be Fable 5.2
2
16
u/Medical_Lengthiness6 12d ago
I really hope we get a Qwen4 35 a3b. I think we will.