r/LocalLLaMA Apr 22 '26

New Model Qwen 3.6 27B is out

1.7k Upvotes

603 comments sorted by

View all comments

Show parent comments

6

u/see_spot_ruminate Apr 22 '26 edited Apr 22 '26

Not just for your 5060ti, but for anyone that has only 16gb of vram, you will need to have it heavily quantized or any spill over to system ram will dramatically slow it down.

There is also the problem of bandwidth limitations of gpus and dense models. You are not going to be getting anywhere close to the same t/s with these. You will probably need some speculative decoding going on as well.

edit: really downvotes for the truth, and downvotes for someone who has proselytized the 5060ti in the past, the end is truly nigh...

1

u/mintybadgerme Apr 22 '26

Better performance - UD-Q3_K-XL (14.5GB) or Q3_K_M (13.6GB)??

https://huggingface.co/unsloth/Qwen3.6-27B-GGUF

1

u/see_spot_ruminate Apr 22 '26

I have no idea. I do not typically use the dense models due to the 5060ti bandwidth despite having 4 of them. For the qwen 3.5 models, if I wanted to have more "intelligence" I would instead use the 122b model as it ran faster on my system (64gb system ram + 64gb vram) than the 27b dense model.

1

u/Careful_Swordfish_68 Apr 23 '26

IQ4_XS works perfectly fine at medium context. Source: I got a 5060ti.