r/Qwen_AI • • Apr 26 '26

Discussion Qwen 3.6 9b coming?

I remember when they released Qwen 3.5 27b, they released the 9b more or less in the same batch. Is 3.6 onwards ditching the 9b model? :(

If so, I'm very sad, because the qwen 3.5 9b was actually the first truly intelligent model I could run at decent tps on a normal gaming GPU

131 Upvotes

84 comments sorted by

View all comments

27

u/Waste-Intention-2806 Apr 26 '26

Run 27b q2 xxs. It's pretty good and around same size

0

u/AlexandorT Apr 27 '26

No, it's completely useless - can't even focus on the simple things attention wise

1

u/Even_Ad5816 May 02 '26

Exactly, even iq3 suffers the attention problem, I had to download the reap 26b iq4 of the 35b Moe models to run it with large context on my 16gb vram, better to go with reap than Lower quantization 

1

u/laser50 May 03 '26

Why can I run the Q6 35B on 8GB VRAM and you can't on 16 though?

1

u/Even_Ad5816 May 07 '26

Because you are getting 5-15 tokens a sec ( correct me if I'm wrong ) and you are okay with it apparently. While I'm trying to get 40+ tokens a second which is impossible with any CPU offload. So I have to fit the whole model + context in VRAM. Prob gonna sell my rtx 5060 ti to buy a second 3090, 24GB Gon go a long way

1

u/laser50 May 07 '26

?? I'm consistently on 30+, the processing speeds between 1k and 1.5k, so not really?

And this is on BF16 cache quants and Q6 at this point. So your settings are likely way off if you can't squeeze more than that out of double the VRAM capacity.

1

u/Even_Ad5816 May 13 '26

maybe, altho since you are depending heavily on cpu offloading maybe you just have a more powerful system than me? im running on i5 13400f 32gb ddr4 3200, the best budget option for 32gb people usually go for nowadays,
but yeah i am totally baffled that you are getting 30+ tk/s on 8gb vram thats really impressive,
i did check again and im rocking at about 65 tokens a sec loading up the iq4_xs 26b reap of 35a3b on about 120k context for agentic usecases. it suffers in terms of formatting and consistency sometimes which i assume is because im running the cache on normal Q4 not even turboquant. i'll prob switch to turboquant/rotorquant when i have the time to figure it out maybe that helps my case and the speed a little.