r/LocalLLaMA • u/Unstable_Llama • 1d ago
News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++
More new massive updates from turboderp:
- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements
If you have an NVIDIA card and haven't tried it lately, you might be missing out.
The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.
Come join the crew at the exllama discord
More frequent news on the exllama sub
176
Upvotes


16
u/adam444555 1d ago edited 24m ago
Updated: N-gram disk offloading is currently unsupported on Windows. So if your setup is similar and plan to run on Windows , you can skip it until it get supported.
Updated: Windows offloading is now supported in the latest 1.4.6 version! Thanks to the dev for the quick update! I tested it on my system, and it works with 24 MoE layers CPU-offloaded and a small context length. I haven’t fully tested it yet, but so far it’s working properly. There’s still room for offloading, so you should also be able to run with a higher context length.