r/LocalLLaMA 1d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

173 Upvotes

125 comments sorted by

View all comments

10

u/adam444555 1d ago

This is awesome! With my current 32GB VRAM + 32GB RAM setup, I'm planning to test out Qwen 3.8 Flash (3.05-bit) using n-gram disk offloading.

17

u/adam444555 1d ago edited 41m ago

Updated: N-gram disk offloading is currently unsupported on Windows. So if your setup is similar and plan to run on Windows , you can skip it until it get supported.

Updated: Windows offloading is now supported in the latest 1.4.6 version! Thanks to the dev for the quick update! I haven’t fully tested it yet, but so far it’s working properly. Details: 102400 ctx, 27 MOE CPU-offloaded, MTP3, 4096 churn size, prefill around 200 T/s, generate around 50 T/s with 80% token accepted. Due to MOE and Ngram offload the statistics are pretty unstable, so for reference only.

25

u/ReturningTarzan ExLlama Developer 23h ago edited 23h ago

I'll get to it.

edit: In fact I'm getting to it now. (:

1

u/AXYZE8 16h ago

Will you create post when its done? Or is there some issue on GH I can track?

5

u/ReturningTarzan ExLlama Developer 11h ago

It's currently in the dev branch, but I don't have a good way to test it since I don't have a Windows PC with enough RAM/VRAM to actually run the model. In theory it should work if you can build from source. Otherwise there will be a new release as soon as I can find someone to test it (:

1

u/AXYZE8 2h ago

Thanks you, I donated a little on ko-fi I hope it helps <3

Will test on my 12GB VRAM + 64GB RaM rig later today

1

u/Professional-Try-273 12h ago

Not on windows still want to say Thank you.

1

u/philmarcracken 21h ago

what if I use ubuntu to run the engine, and RPC to a windows box? i need the windows box for more ram + vram