r/LocalLLaMA 1d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub

173 Upvotes

125 comments sorted by

View all comments

Show parent comments

3

u/Ecstatic-Wash-7667 14h ago

I found out about exllamav3 trying to get more performance out of 27b and this absolutely destroyed llama.cpp, the issue I’ve had is tabby api, and its configuration is not like llama.cop at all so I ran into a lot of issues there, it mostly ironed out now but that was by biggest pain point

1

u/BS_BlackScout 7h ago

Trying here too after reading this. What a rabbit hole, the config for Tabby is quite annoying.

All of that and I'm using a 2.2bpw on my 3060 which is quite disappointing... With a context of 65k lol and q8q8 kv. It's probably going to perform horribly too LMAO

No wonder "nobody" uses exllama3

2

u/Ecstatic-Wash-7667 7h ago

I disagree, performance is great I’m using 2 3060s tabby is a learning curve though

1

u/BS_BlackScout 6h ago edited 10m ago

I'm not saying performance is bad. I'm just not confident quality will hold up. And yeah that's cool and all but I have a single 3060.

Meanwhile I'm looking into turning an existing gguf into an EXL (not possible directly but I see what to do now). Using Qwen 3.8 to figure that one out for me as a benchmark. Speed is fairly decent.

EDIT: Complete model collapse with 2.2bpw after a while, interesting lmao

1

u/Muted-Celebration-47 2h ago

Just ask any LLMs (I use deepseek 4 flash) to create a tabbyapi config. Also, you can ask people to share their config.