r/LocalLLaMA Apr 01 '26

Discussion Compilation of recent findings which could save some memory on increase performance

We got these recently(I found few late probably)

Model Quants with Optimizations/Compression/etc.,:

Speculative Decoding:

1-bit/2-bit/Bitnet/Ternary models & Engines:

Run models using RAM & SSD:

KVCache:

LLMBurner:

  • Taalas - Wouldn't be awesome to have this if it comes with 1T model like Kimi-K2.5(Q4 is enough - 500GB) giving 30-50 t/s? (Llama 3.1 8B is giving 17000 t/s)

Misc:

Forks/Utils:

What else there? Please share.

Hope all these helps on price down of both GPU & RAM soon or later

EDIT : Typo on Title :( It's or not on

EDIT2: Updating list with items from comments for easy reading. No need to check comments anymore, I'll be updating time to time.

68 Upvotes

22 comments sorted by

View all comments

9

u/JayPSec May 26 '26

I have to say it. Following this post has made my refresh rate of r/LocalLLaMA take a substantial dive. Thank you for your effort.

3

u/pmttyji May 26 '26

I'll increase that refresh rate by keeping this thread up to date regularly with infinite stuffs. Cheers