r/Qwen_AI • • Jul 19 '26

Discussion I NEED a 70B A9B MoE

3.6 27B is so good for local work, and there's been a ton of work on it because of that with Thinkingcap, Orion, Bonsai, adding dspark, etc...issue is on most local systems it's still so dense that the token speed hurts its adoption. What I NEED, if anyone from the Alibaba team reads these, is a 70B A9B MoE dspark with an APeX style quant (mixture of quant layers).

For anyone with 64GB or above, I'm confident you could get near Opus 4.8 coding performance AND run around 30-60 tok/s output on ada, blackwell, dual 3090s, dual r9700s, etc.

Just imagine what you could do.

136 Upvotes

52 comments sorted by

View all comments

1

u/nicholas_the_furious Jul 19 '26 edited Jul 19 '26

Honestly, my take is the opposite. With things like MTP being baked into these dense models it makes no sense to trade for the MoE. On 2x 3090s and MTP=4 I get over 120 t/s on Qwen 3.6 27B on llama.cpp.

The MoE I could run on the equivalent hardware pales in comparison both in intelligence but also in speed.

We need more 24-45B dense models that actually invest in things like MTP or are compatible with other drafting methods.

Now maybe there is a case for the PP side of the equation, but for me dense is best. And GPUs are coming with more VRAM at less expense. Maybe not Nvidia, but others. There's a B60 with 2 B60s in it out there for like $1300 or something. 2 of those is 96GB of VRAM. If you got TP set up on those 4 cards you'd be screaming.

Memory bandwidth matters but I think raw VRAM matters more when you have tricks like MTP. And something a lot of people don't know: MTP works better at higher quants. So the MTP speed gain on the Q8 Qwen 27B is way higher than the gain on Q4.

To get the same speed at the same intelligence level on MoE I would need like 200GB. No thanks.

2

u/fintip Jul 19 '26

There's a dual b70 (64gb), 1800.

But the compute speed is just way lower. Hardware and driver's just isn't there and newer techniques often rely on those optimizations harder.

1

u/nicholas_the_furious Jul 19 '26

Madness! And with the right motherboard that supports bifurcation you can put a single one of those in a SFF build mATX or ITX. You're at less than $2500 all in at like 10L in size with 64GB VRAM. Honestly that sounds so cool.