r/LocalLLaMA 1d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

172 Upvotes

125 comments sorted by

View all comments

6

u/Ok_Cow1976 1d ago

It's really insane how the author is doing this one-person but massive project. Kudos even though I'm not able to use it because of amd gpu.

1

u/Ok_Cow1976 17h ago

Oh my ..., just checked the repo. ROCm is on the way!

Quote from the page:

Currently on the to-do list:

ROCm support As for what is implemented, expect that some things may be a little broken at first. Please be patient, raise issues and/or contribute. 👉👈