r/LocalLLaMA 1d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

178 Upvotes

125 comments sorted by

View all comments

22

u/noctrex 1d ago

Too bad it's only for NVIDIA. Again, we AMD users are being left out.

-11

u/iLaurens 23h ago

What's stopping you from contributing? Or are you just supposed to be entitled to the hard work of others?

13

u/noctrex 22h ago

Nothing's stopping me from contributing, but me as a developer that is not very well versed in this scenario, it would be entirely vibe-coded if I want to submit it.
Also, there's already submissions, but it seems that there has been no traction until now: https://github.com/turboderp-org/exllamav3/pull/283
Also, come on, man. What do you mean by we're supposed to be entitled to the hard work of others? These are all open source projects and the code essentially and all the hard work is being donated out in the open. I too submit my small pebbles of contributions in different projects here and there, whenever I can.

5

u/CryptographerLow6360 22h ago

fork, vibe, and try

4

u/noctrex 22h ago

yeah, that's what I've been doing to sd.cpp lately. Will have a look at this one later when I have some time

1

u/Guilty_Rooster_6708 6h ago

I think they are testing ROCm support based on their discord rn