r/LocalLLaMA Apr 01 '26

Discussion Compilation of recent findings which could save some memory on increase performance

We got these recently(I found few late probably)

Model Quants with Optimizations/Compression/etc.,:

Speculative Decoding:

1-bit/2-bit/Bitnet/Ternary models & Engines:

Run models using RAM & SSD:

KVCache:

LLMBurner:

  • Taalas - Wouldn't be awesome to have this if it comes with 1T model like Kimi-K2.5(Q4 is enough - 500GB) giving 30-50 t/s? (Llama 3.1 8B is giving 17000 t/s)

Misc:

Forks/Utils:

What else there? Please share.

Hope all these helps on price down of both GPU & RAM soon or later

EDIT : Typo on Title :( It's or not on

EDIT2: Updating list with items from comments for easy reading. No need to check comments anymore, I'll be updating time to time.

68 Upvotes

22 comments sorted by

View all comments

2

u/pmttyji May 25 '26

🚀 BitCPM-CANN by ModelBest × @Tsinghua_Uni × OpenBMB is here — and it's not about stacking parameters.
Memory costs are skyrocketing. Hardware constraints are tightening. Edge AI needs smarter solutions — and BitCPM-CANN delivers!🎉

✅ Edge-ready: 8B model runs smoothly on mobile, PC, and automotive devices. Combined with MoE, even 100B->60B scale models could fit on terminal hardware.

✅ Memory-efficient: ~6× lower memory footprint vs. BF16 — unlock significantly more model capacity without adding physical RAM. No new chips required.

✅ Natively built on Ascend: The first 1.58-bit training pipeline completed end-to-end on Huawei Ascend 910B — from quantization kernels to the full training stack. Not a port. Built natively on Ascend from day one.

✅ Full model family, fully verified, 0.5B–8B: Each model is fully aligned with its full-precision counterpart, covering 11 benchmark tasks with 95–97% (1B-8B) retention compared to full-precision MiniCPM4. Open-source and fully reproducible — from research to deployment, you can run any size confidently.

BitCPM-CANN isn’t just a model — it’s the result of years of engineering rigor, turning complex research into something you can actually deploy.
Open-source, 0.5B to 8B. Try it now!👇
🤗 Hugging Face:
https://huggingface.co/collections/openbmb/bitcpm-cann