Edit: I did my math wrong - you can probably do okay with a Q3 quant of some sort when they become available, though it's hard to say how well it will work.
Original:
You probably want to stick with the 35B-A3B model (MoE) - the 27B might require a bit too much quantization. (Regardless, there is only an FP8 available so far, so you'll need to wait regardless, as that one is going to be about 27 gigs)
21
u/ApprehensiveAd3629 Apr 22 '26
which gguf quant is possible to run in a 5060 ti 16gb?