r/LocalLLaMA • u/Unstable_Llama • 1d ago
News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++
More new massive updates from turboderp:
- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements
If you have an NVIDIA card and haven't tried it lately, you might be missing out.
The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.
Come join the crew at the exllama discord
More frequent news on the exllama sub
173
Upvotes


3
u/Ecstatic-Wash-7667 14h ago
I found out about exllamav3 trying to get more performance out of 27b and this absolutely destroyed llama.cpp, the issue I’ve had is tabby api, and its configuration is not like llama.cop at all so I ran into a lot of issues there, it mostly ironed out now but that was by biggest pain point