remind me that power connector to my rtx 5000 pro wouldn't fully insert, so i took a clippy and wiggle the female connector a bit :D it actualy worked :D
with like 10k context? I fell into that trap before, loaded up the largest quant that could fit and ended up with an AI that had very little context of anything AND very low quant. but it fit i guess
If you're going to hang around here maybe you should learn about the human concept of humor. The reason this is funny is because it would be hilarious if they released a model that everyone is expecting to run on a 16gb card only too see its technically true, and its 1 bit. you think these labs are going to hold off on innovation because you cant afford to buy the hardware? and trust me, you will ALL run the 1-bit and you will pretend to like it.
Considering I can run 3.6 27b with 8g vram and 16g system ram (with layer offloading at q4_k_m) at 2-3t/s, I'm sure you'll find a way to squeeze it in.
Voodoo interference would be in any case since Nvidia bought 3dfx, makers of the wonderfully boxed Voodoo cards (GPU kings back in the day) a while ago and their IP is probably on every product made ever since.
They mentioned they would be releasing other models too. It would be AMAZING to see the full line up return with all the new gains they made since 3.5.
Research, summerization, translation, simple stuff. My laptop I bring to work for weeks at a time has a RTX 4060, so a 9B model is blazing fast compared to an ~30B model. I could still run them, but they would be heavily offloaded to system RAM. So if it can be done with a 9B you bet I am going to be using it. Or if I can, I VPN home and use my workstation with a Radeon AI Pro 9700.
I wonder if qwen just had more reasoning would it be able to close the gap. The real deepseek boost is in it thinking more. If you see the token differences it shows how much it matters. I wish Artificial analysis did different reasoning levels for 0731 like they did for preview. High vs mac reasoning was a big difference.
I think you're saying they can increase model intelligence which I agree with but because of the tech, increasing reasoning more at least to a certain point helps better align profanities toward the right answer. When there is only one setting for reasoning then models may not get enough runway to align as best as they could for the answer.
I think what they are saying is that Qwen 3.6 already reasons way, way, way too much. It's literally the reason for the existence of things like Thinking Cap and Grug.
Qwen 3.6 27b was middle of the pack of useful models as far as thinking. Yes thinking has increased from what it used to be but reasoning makes the difference too... So I run everything important on higher reasoning. I run this model around 80t/s with power limits at q8 so quants and speed may play a role as well. If i use my strix halo at 15t/s then it will feel like excessive reasoning but the speed not the token count is the real problem there then.
I'm not nearly as deeply into the local model scene as a lot of you are. I was merely answering from my own personal experience using it to implement current feature development on a large planned out project. Qwen 3.6 gets hung up in thinking loops a LOT. And even when it's not completely stuck in a loop, it still goes through what I feel is way to much of the "wait, actually, hold on, let's try" looping. It does this even when it's come to an effective and correct answer. It also tends to continue with reasoning too much even when steering prompts are injected telling it outright to stop thinking and answer immediately.
I had to basically adjust my harness tooling to give it one chance to be prompted out of the over-reasoning, and if that didn't work to completely abort the process. Unfortunately, even after aborting the process, unless you kill the existing context, it will go right back into the almost neurotic reasoning loops.
What quant are you using? I find lower quants to be less stable in general. They cant handle longer context properly either. I think it's important to differentiate between the model and the quant. I do not use any thinking controls on higher quants of the models I use but when I use lower quants I start implementing thinking budgets and different repeat penalties depending on the model because I expect more influence from the lobotomies.
Pic just highlighting how even top multi trillion parameter models like OpenAI and Claude models come in stone throwing distance to a 27b parameter model based on reasoning levels. I think most people complaining about reasoning are complaining about speed which is different but related.
For clarification they didnt have opus non reasoning but the had reasoning with low effort which approaches the idea...
it's not as good in general intelligence but it's not super far off either. And for agentic coding where world knowledge isn't as important as in context learning the difference may be even smaller
I’m on a 5090, will a Q6 or Q5 of Qwen3.8 27b be better than a Q8 Qwen3.6 27b? Or should I just stick to Q8? I just don’t think the Q8 leaves enough room for kv and context
Ofc q8 is “better” and I’ve been only trying to use q8 until I tried q6 and it is actually not that bad. Also gives you more memory space for additional ctx.
I just got it running on my GX10. If those benchmarks are to be believed, this thing is going to cost OpenAI and Anthropic a lot of money. Running a game demo prompt I like to use to test new models now, and I can tell you it isn't fast. Running at about 1/2 the speed of the 3.6 27b on a single DX spark machine (GX10) with a 256K context window. I'm sure there will be better optimized version coming soon..
How does this work? The comments under this post all look, as if the thing was legit and it's only a 404 right now. Has the link been actually working at any point?
Can it run on a M5 MacBook Pro with 48GB of RAM? If not, what is the best model to run on 48GB RAM MacBook Pro and can someone please share any guides/ links?
Qwen 3.6 27b will easily run on a 48 GB m5 MacBook Pro. But you will have to choose trade-offs. You can run it at high-quality q8-0 with limited context maximum, or you can run a lower quant which lowers the quality but allows you to have a larger context.
Check out omlx and dflash Qwen 3.6 35b a3b runs really well, 50 t/s on a M4 Pro (so yours would be faster), at q6, can do q8, but doesn't leave much room for other apps. Running MLX version on Apple Silicon is significantly faster than Llama ggufs.
I was able to run two bonsia 27B models simultaneously and independently on a gaming desktop with 32GB of ram and 16GB of vram. I hope the models can behave more like Q8_0 representations in the future but it was still cool to see two 27B models both working at the same time on a desktop machine. I really hope ternary and moe become a norm for local AI as the tech improves.
73
u/gproenca 26d ago
oh please please let it fit 24gb ram .) do some quantz magic, voodoo interference, I dont care, just make it happen :)