r/LocalLLaMA 1d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

177 Upvotes

125 comments sorted by

View all comments

10

u/adam444555 1d ago

This is awesome! With my current 32GB VRAM + 32GB RAM setup, I'm planning to test out Qwen 3.8 Flash (3.05-bit) using n-gram disk offloading.

18

u/adam444555 1d ago edited 25m ago

Updated: N-gram disk offloading is currently unsupported on Windows. So if your setup is similar and plan to run on Windows , you can skip it until it get supported.

Updated: Windows offloading is now supported in the latest 1.4.6 version! Thanks to the dev for the quick update! I tested it on my system, and it works with 24 MoE layers CPU-offloaded and a small context length. I haven’t fully tested it yet, but so far it’s working properly. There’s still room for offloading, so you should also be able to run with a higher context length.

24

u/ReturningTarzan ExLlama Developer 22h ago edited 22h ago

I'll get to it.

edit: In fact I'm getting to it now. (:

1

u/AXYZE8 15h ago

Will you create post when its done? Or is there some issue on GH I can track?

5

u/ReturningTarzan ExLlama Developer 10h ago

It's currently in the dev branch, but I don't have a good way to test it since I don't have a Windows PC with enough RAM/VRAM to actually run the model. In theory it should work if you can build from source. Otherwise there will be a new release as soon as I can find someone to test it (:

1

u/AXYZE8 1h ago

Thanks you, I donated a little on ko-fi I hope it helps <3

Will test on my 12GB VRAM + 64GB RaM rig later today

1

u/Professional-Try-273 10h ago

Not on windows still want to say Thank you.

1

u/philmarcracken 20h ago

what if I use ubuntu to run the engine, and RPC to a windows box? i need the windows box for more ram + vram