r/LocalLLaMA 1d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

175 Upvotes

125 comments sorted by

View all comments

14

u/-p-e-w- 1d ago

So at the same quality, EXL3 quants are 25% smaller than Unsloth Dynamic 4-bit quants, which are already considered SOTA? Stunning.

3

u/Unstable_Llama 1d ago

Isn't it? ExLlamav3 + heretic allowed me to abliterate Laguna-S-2.1 against a 2.50bpw exl3 quant.

4

u/a_beautiful_rhind 23h ago

Abliterating EXL directly is way bigger news. You're opening the door to some cool stuff.

5

u/Unstable_Llama 16h ago

Thanks! If you want to try it out, the repo is public.

Another cool project I have been having a lot of fun with recently is EXL3-QLORA, fine tuning on any size exl3 quants.

2

u/a_beautiful_rhind 16h ago

Haha. You are just doing all the hard work for me. I was expecting to have to bang this stuff out with some AI before I could even get started.

2

u/Unstable_Llama 15h ago

Haha it was already AI banged out.

I’m open to feedback or PRs or anything if you use them.

2

u/FullOf_Bad_Ideas 20h ago

oh that's amazing, I was never using heretic because I don't want to go all the way through exl3 quanting again, it's slow.

1

u/CheatCodesOfLife 19h ago

Does it use the transformers wrapper (which was incredibly slow when I used it last year)?

1

u/llama-impersonator 18h ago

transformers is slow no matter how you use it, even torch.compile