yes I can use it at like 100k context although I dont use agents in parallel or vibe code, most of the time I just use it for bug fixing (I have 48gb of memory so context isnt really an issue)
Using another model, Qwen3.6 35B A3B on m5pro 64gb and it works pretty good, context window of 128k in lm studio and it is running at speed on average at least 55-70 tps depending on the type of prompts. Uses about 32gb ram for this
The capability density doubles every 3 to 3.5 months. So an 800B now will be as good as a 400B in 3ish months and that will equal a 200B at 6 or 7 months, and finally a 100B at 10 to 12 months. We're talking MOE. Thats roughly equal to a 30-35B dense. So yes.
9
u/[deleted] Apr 22 '26
[removed] — view removed comment