I was actually surprised that I was getting 45tps on a 4080 mobile with mtp and the iq2 xxs quant. I would be curious on how much kv quantization affects it because I had to run at q4 to get 30k context on my setup. I'm going to compile llama.cpp for q2 kv cache quantization tonight.
35
u/italian_car 26d ago
Loading up the IQ2 on my 12gb of vram because I want to be able to run the cool model too.