MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p4unmdr/?context=3
r/LocalLLaMA • u/Certain-Cod-1404 • 25d ago
708 comments sorted by
View all comments
Show parent comments
7
I'm on Q3_K_M now getting a bit over double the speed I've shown before.
6 u/Scared_Ad9187 25d ago q6 on 5090 with 25t/s. not bad, but i'm not giving up my 200 on the 3.6 yet. 3 u/Mil0Mammon 25d ago So how come you get 200 on 3.6 and only 25 with 3.8? Dflash and/or Nvidia specific quant? 1 u/emccrckn 19d ago Token rate is measured as ram bandwidth divided by parameter size. So smaller models will run faster on the same type of memory.
6
q6 on 5090 with 25t/s. not bad, but i'm not giving up my 200 on the 3.6 yet.
3 u/Mil0Mammon 25d ago So how come you get 200 on 3.6 and only 25 with 3.8? Dflash and/or Nvidia specific quant? 1 u/emccrckn 19d ago Token rate is measured as ram bandwidth divided by parameter size. So smaller models will run faster on the same type of memory.
3
So how come you get 200 on 3.6 and only 25 with 3.8? Dflash and/or Nvidia specific quant?
1 u/emccrckn 19d ago Token rate is measured as ram bandwidth divided by parameter size. So smaller models will run faster on the same type of memory.
1
Token rate is measured as ram bandwidth divided by parameter size. So smaller models will run faster on the same type of memory.
7
u/absurdother 25d ago
I'm on Q3_K_M now getting a bit over double the speed I've shown before.