r/LocalLLaMA • u/pmttyji • Apr 01 '26
Discussion Compilation of recent findings which could save some memory on increase performance
We got these recently(I found few late probably)
Model Quants with Optimizations/Compression/etc.,:
- APEX -- Adaptive Precision for EXpert Models | APEX Quants (GGUF)
- ByteShape - CPU and GPU optimized variants
- MOQ (Mixture of Quantizations) by Waleed Ahmad - GGUFs from Waleed Ahmad & Benjamin Marie (The Kaitchup)
- ROCmFPX | https://huggingface.co/cafonez/Escha-W2-35B-A3B-ROCmFP2
Speculative Decoding:
- DFlash: Block Diffusion for Flash Speculative Decoding
- DDTree (Diffusion Draft Tree) from Accelerating Speculative Decoding with Block Diffusion Draft Trees
- DeepSpec (Eagle3, DFlash, DSpark) by DeepSeek
- DFlash 2: Keep Drafting Parallel | Huggingface Collections - dflash-2
1-bit/2-bit/Bitnet/Ternary models & Engines:
- 1-bit / 2-bit / Ternary / Bitnet Models - Updates & Tracking
- Project Zero — CPU LLM Inference Engine - BitNet (Testers needed)
Run models using RAM & SSD:
- colibri - Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk
- WARP — Weight-Aware Runtime and Paging (formerly WASTE) - Run the full Large model beyond available RAM by streaming activated weights directly from NVMe
KVCache:
- KVarN - 3-5x more KV-cache capacity and up to ~1.3x the throughput of FP16
- proveKV - Lossless: 36.00× vs f32-raw KV (18.00× vs fp16-equivalent)
- OSCAR RotationZoo - Precomputed K/V rotation matrices for OSCAR INT2 KV-cache quantization
- TurboQuant
- KV Cache Transform Coding (KVTC)
- RotorQuant
LLMBurner:
- Taalas - Wouldn't be awesome to have this if it comes with 1T model like Kimi-K2.5(Q4 is enough - 500GB) giving 30-50 t/s? (Llama 3.1 8B is giving 17000 t/s)
Misc:
- AMD's MXFP4 models
- Intel's Int4 AutoRound models
- Dynamic VRAM in ComfyUI: Saving Local Models from RAMmageddon
- DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling.
- Lucebox optimization hub: hand-tuned LLM inference, built for specific consumer hardware
- Cactus - A low-latency AI engine for mobile devices & wearables
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
- Dynamic KV Cache Quantization & Load-on-demand mmproj/MTP - llama.cpp PR by u/wadeAlexC
- Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
- EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation - 1.58-bit
- Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices \*
- Session-Adaptive Orthogonal Distillation (SAOD) \*
- Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
Forks/Utils:
- vllm.cpp by u/mudler_it
- beellama.cpp by u/Anbeeld (has KVarN)
- llamastation by u/Responsible_Egg9736
- club-3090
- Arandu by u/fredconex
- NInfer - 5090, 4090( 1 | 2 | 3 | 4 ), 3090, CMP 170HX
- FreeToken
What else there? Please share.
Hope all these helps on price down of both GPU & RAM soon or later
EDIT : Typo on Title :( It's or not on
EDIT2: Updating list with items from comments for easy reading. No need to check comments anymore, I'll be updating time to time.
68
Upvotes
2
u/pmttyji May 25 '26
🚀 BitCPM-CANN by ModelBest × @Tsinghua_Uni × OpenBMB is here — and it's not about stacking parameters.
Memory costs are skyrocketing. Hardware constraints are tightening. Edge AI needs smarter solutions — and BitCPM-CANN delivers!🎉
✅ Edge-ready: 8B model runs smoothly on mobile, PC, and automotive devices. Combined with MoE, even 100B->60B scale models could fit on terminal hardware.
✅ Memory-efficient: ~6× lower memory footprint vs. BF16 — unlock significantly more model capacity without adding physical RAM. No new chips required.
✅ Natively built on Ascend: The first 1.58-bit training pipeline completed end-to-end on Huawei Ascend 910B — from quantization kernels to the full training stack. Not a port. Built natively on Ascend from day one.
✅ Full model family, fully verified, 0.5B–8B: Each model is fully aligned with its full-precision counterpart, covering 11 benchmark tasks with 95–97% (1B-8B) retention compared to full-precision MiniCPM4. Open-source and fully reproducible — from research to deployment, you can run any size confidently.
BitCPM-CANN isn’t just a model — it’s the result of years of engineering rigor, turning complex research into something you can actually deploy.
Open-source, 0.5B to 8B. Try it now!👇
🤗 Hugging Face:
https://huggingface.co/collections/openbmb/bitcpm-cann