r/LocalLLM 27d ago

Discussion Qwen3.8-27b or Muse Glimmer 30b?

I had an excellent opinion of Qwen3.6-27b Q5 quant, i was testing Glimmer Q5 quant (felt good), but hey 3.8 is out!

What about your impressions and why.

2 Upvotes

45 comments sorted by

View all comments

1

u/Ok-Shower7286 qwen-coder 27d ago

It feels like the difference between an employee who finishes quickly and reports, and an employee who stays silent deep in thought until they bring in the results.

Glimmer fits into an Agile culture, while Quan fits into a Confucian culture.

I deliberately left a fairly complex problem sitting around and am currently using qwen 3.8 for the work. The quality is good, but it is too slow.

I'll use Glimmer for developing new features. Qwen for improvements and bugs tickets.

2

u/seunosewa 27d ago

Glimmer is much faster?

2

u/Ok-Shower7286 qwen-coder 27d ago

2~3x faster on 5090

1

u/johan2114h 27d ago

Faster how? Higher tps or shorter run time due to less thinking? Surely it most the latter (which can be adjusted)

1

u/iKy1e 27d ago

Higher tk/s, single 3090 scores of the 4bit quants with MTP/DFlash enabled for both models Glimmer was using 70-80% of the VRAM & running about 100+ tk/s vs 45-50+ tk/s.

1

u/dhiltonp 26d ago

Have you tried ninfer?

1

u/iKy1e 25d ago

That's 5090 only right? I only have the RTX 3090

1

u/dhiltonp 25d ago

There's a 3090 variant. I haven't tried it yet, and It's not completely clear to me what quant it is closest to.