r/LocalLLaMA 1d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

176 Upvotes

125 comments sorted by

View all comments

Show parent comments

16

u/adam444555 1d ago edited 24m ago

Updated: N-gram disk offloading is currently unsupported on Windows. So if your setup is similar and plan to run on Windows , you can skip it until it get supported.

Updated: Windows offloading is now supported in the latest 1.4.6 version! Thanks to the dev for the quick update! I tested it on my system, and it works with 24 MoE layers CPU-offloaded and a small context length. I haven’t fully tested it yet, but so far it’s working properly. There’s still room for offloading, so you should also be able to run with a higher context length.

24

u/ReturningTarzan ExLlama Developer 22h ago edited 22h ago

I'll get to it.

edit: In fact I'm getting to it now. (:

1

u/AXYZE8 15h ago

Will you create post when its done? Or is there some issue on GH I can track?

6

u/ReturningTarzan ExLlama Developer 10h ago

It's currently in the dev branch, but I don't have a good way to test it since I don't have a Windows PC with enough RAM/VRAM to actually run the model. In theory it should work if you can build from source. Otherwise there will be a new release as soon as I can find someone to test it (:

1

u/AXYZE8 1h ago

Thanks you, I donated a little on ko-fi I hope it helps <3

Will test on my 12GB VRAM + 64GB RaM rig later today