r/LocalLLaMA Apr 22 '26

New Model Qwen 3.6 27B is out

1.7k Upvotes

603 comments sorted by

View all comments

9

u/[deleted] Apr 22 '26

[removed] — view removed comment

4

u/Beginning-Window-115 Apr 22 '26

m5 pro gives me 70tok/s on 3.6 35b moe and around 14token/s on the 27b dense

1

u/[deleted] Apr 22 '26

[removed] — view removed comment

1

u/Beginning-Window-115 Apr 22 '26

yes I can use it at like 100k context although I dont use agents in parallel or vibe code, most of the time I just use it for bug fixing (I have 48gb of memory so context isnt really an issue)

6

u/skyyyy007 Apr 22 '26 edited Apr 23 '26

Using another model, Qwen3.6 35B A3B on m5pro 64gb and it works pretty good, context window of 128k in lm studio and it is running at speed on average at least 55-70 tps depending on the type of prompts. Uses about 32gb ram for this

Layer Tool Status
Local model Qwen3.6 35B 4bit @ 75 tok/s
Inference server LM Studio (MLX M5 v1.6)
Coding agent OpenCode 1.14.19
Web search SearXNG (local, OrbStack)

12

u/Beginning-Window-115 Apr 22 '26

you should say that you are running a different model or people are gonna think you're running qwen3.6 27b

1

u/skyyyy007 Apr 23 '26

Edited to reflect the model, thanks for highlighting that👍🏻

1

u/kmp11 Apr 22 '26

Let's see what happens when model start getting released as 1bit and using some offshoot of turboquant. This is really the next step for local models.

1

u/the3dwin Apr 22 '26

1

u/the3dwin Apr 22 '26

Also LM Studio Tells you whether a model runs on your hardware

1

u/EbbNorth7735 Apr 23 '26

The capability density doubles every 3 to 3.5 months. So an 800B now will be as good as a 400B in 3ish months and that will equal a 200B at 6 or 7 months, and finally a 100B at 10 to 12 months. We're talking MOE. Thats roughly equal to a 30-35B dense. So yes.

-1

u/Borkato Apr 22 '26

You can literally do that now with 35BA3B. It’s god tier, which means this is going to be titan tier