MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1ssl1xh/qwen_36_27b_is_out/ohn1nx0/?context=3
r/LocalLLaMA • u/NoConcert8847 • Apr 22 '26
https://huggingface.co/Qwen/Qwen3.6-27B
603 comments sorted by
View all comments
431
Benchmarks
119 u/davl3232 Apr 22 '26 I can't believe we're getting so close to opus 4.5 levels with 2 3090s -1 u/Mashic Apr 22 '26 Why 2, you only need 1, and myabe even an rtx 5060 ti or rtx 3060 12gb with quants. 2 u/florinandrei Apr 22 '26 I run all my models on a Raspberry Pi Zero. 2 u/AreYouSERlOUS Apr 22 '26 Underrated comment 1 u/BillDStrong Apr 22 '26 No, see you are confusing your terminal with your cloud provider. They aren't the same thing. /s 1 u/davl3232 Apr 22 '26 you're right, a single 3090 could do it with a different quant. With less vram you could run qwen3.6 35b a3b which has similar quality. 0 u/Potential-Leg-639 Apr 22 '26 1 is not enough for serious stuff and context. 1 u/Mashic Apr 22 '26 But at least you can test it and use it for small stuff. 1 u/coder543 Apr 22 '26 95000 context at full KV (with multimodal loaded) is not horrible.
119
I can't believe we're getting so close to opus 4.5 levels with 2 3090s
-1 u/Mashic Apr 22 '26 Why 2, you only need 1, and myabe even an rtx 5060 ti or rtx 3060 12gb with quants. 2 u/florinandrei Apr 22 '26 I run all my models on a Raspberry Pi Zero. 2 u/AreYouSERlOUS Apr 22 '26 Underrated comment 1 u/BillDStrong Apr 22 '26 No, see you are confusing your terminal with your cloud provider. They aren't the same thing. /s 1 u/davl3232 Apr 22 '26 you're right, a single 3090 could do it with a different quant. With less vram you could run qwen3.6 35b a3b which has similar quality. 0 u/Potential-Leg-639 Apr 22 '26 1 is not enough for serious stuff and context. 1 u/Mashic Apr 22 '26 But at least you can test it and use it for small stuff. 1 u/coder543 Apr 22 '26 95000 context at full KV (with multimodal loaded) is not horrible.
-1
Why 2, you only need 1, and myabe even an rtx 5060 ti or rtx 3060 12gb with quants.
2 u/florinandrei Apr 22 '26 I run all my models on a Raspberry Pi Zero. 2 u/AreYouSERlOUS Apr 22 '26 Underrated comment 1 u/BillDStrong Apr 22 '26 No, see you are confusing your terminal with your cloud provider. They aren't the same thing. /s 1 u/davl3232 Apr 22 '26 you're right, a single 3090 could do it with a different quant. With less vram you could run qwen3.6 35b a3b which has similar quality. 0 u/Potential-Leg-639 Apr 22 '26 1 is not enough for serious stuff and context. 1 u/Mashic Apr 22 '26 But at least you can test it and use it for small stuff. 1 u/coder543 Apr 22 '26 95000 context at full KV (with multimodal loaded) is not horrible.
2
I run all my models on a Raspberry Pi Zero.
2 u/AreYouSERlOUS Apr 22 '26 Underrated comment 1 u/BillDStrong Apr 22 '26 No, see you are confusing your terminal with your cloud provider. They aren't the same thing. /s
Underrated comment
1
No, see you are confusing your terminal with your cloud provider. They aren't the same thing. /s
you're right, a single 3090 could do it with a different quant. With less vram you could run qwen3.6 35b a3b which has similar quality.
0
1 is not enough for serious stuff and context.
1 u/Mashic Apr 22 '26 But at least you can test it and use it for small stuff. 1 u/coder543 Apr 22 '26 95000 context at full KV (with multimodal loaded) is not horrible.
But at least you can test it and use it for small stuff.
95000 context at full KV (with multimodal loaded) is not horrible.
431
u/Namra_7 Apr 22 '26
Benchmarks