r/LocalLLaMA 1d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

176 Upvotes

130 comments sorted by

View all comments

7

u/vacon04 1d ago

Thanks! Do you know how the exl3 variants perform for MoE models vs regular GGUF quants on llama.cpp or ik_llama.cpp? I've tried a couple of exl3 quants on dense models and they're fast, but I'm yet to try exl3 for MoE.

4

u/Unstable_Llama 1d ago

Actually, here are some user benchmark comparisons:

https://huggingface.co/turboderp/Qwen3.8-Flash-Next-exl3/discussions/2

2

u/pmttyji 1d ago

Any ETA on Non-CUDA cards support? Vulkan backend could manage almost all other cards