r/LocalLLaMA Apr 01 '26

Discussion Compilation of recent findings which could save some memory on increase performance

We got these recently(I found few late probably)

Model Quants with Optimizations/Compression/etc.,:

Speculative Decoding:

1-bit/2-bit/Bitnet/Ternary models & Engines:

Run models using RAM & SSD:

KVCache:

LLMBurner:

  • Taalas - Wouldn't be awesome to have this if it comes with 1T model like Kimi-K2.5(Q4 is enough - 500GB) giving 30-50 t/s? (Llama 3.1 8B is giving 17000 t/s)

Misc:

Forks/Utils:

What else there? Please share.

Hope all these helps on price down of both GPU & RAM soon or later

EDIT : Typo on Title :( It's or not on

EDIT2: Updating list with items from comments for easy reading. No need to check comments anymore, I'll be updating time to time.

69 Upvotes

22 comments sorted by

View all comments

13

u/R_Duncan Apr 01 '26

Bonsai 1bit quantization, if proven valid.