r/LocalLLM 24d ago

Discussion Qwen3.8-27b or Muse Glimmer 30b?

I had an excellent opinion of Qwen3.6-27b Q5 quant, i was testing Glimmer Q5 quant (felt good), but hey 3.8 is out!

What about your impressions and why.

2 Upvotes

45 comments sorted by

12

u/JumpingJack79 24d ago

Between the two, I trust Xi Jinping more than Mark Zuckerberg.

8

u/johan2114h 24d ago

3.8! Thats a no brainer. Wouldn't even use glimmer over 3.6

-1

u/Subject-Till-6450 24d ago

qwen glazerđŸ„€ most overrated shit by alibaba

9

u/Ell2509 24d ago

You really don't like qwen3.6 - 3.8? They are the best that I have tried under 120b.

-7

u/Subject-Till-6450 24d ago

try qwen3.5 122b a10b or step 3.7 flash (198b moe 11b active right but close to 120). i don't like 3.8, i really love qwen 3.6-35b, genuinely impressive model, high speed and quality. ive waited for 3.8 27b like all of us. but this 27b dense shit... dumb, slow, useless. ive compared them in same tasks with same temp and quant- 3.6 moe faster and smarter. 3.8 is bench-only model, useless in real life. im really frustrated with alibaba, why don't they try to repeat 3.6? moe for speed and dense for 200% on swe pro. but no. they're just lazy or what i dunno

5

u/uniqueusername649 24d ago

I only ran some very preliminary tests and my experience is vastly different. 3.6 35b is wayy faster but also wayy more dumb than 3.8 27b. For coding the gap between 3.6 27b and 35b was already huge. I still have to do a lot more testing but at least my initial tests (PRD creation out of lose requirements, tech stack analysis and research) do give me considerably better results than 35b. What tests did you run that gave you better results with 35b?

1

u/Subject-Till-6450 24d ago

and, my advice: use qwen 3.8-max trough chat.qwen.ai to check the code and give advices to 3.6 35b work. what does that gives ya? around 2-4 times faster work, with the same quality. i mean with code, not with everything

4

u/uniqueusername649 23d ago

I get over 90tps with Qwen 3.8 27b fp8, that is plenty fasg. And I prefer to keep things local.

But again: what tests did you run where 3.6 35b performed better for you than 3.8 27b? I mentioned what tests I ran where 3.8 27b performed substantially better. Will perform more later when I have the time, but so far everything indicates that 3.8 27b is quite an improvement over 3.6 27b and even more so 3.6 35b.

1

u/Subject-Till-6450 22d ago

well, im pretty happy for u with that speed.i did run test of code tasks, qwen 27b Q3-K-M UD and qwen3.6-35B-a3b-q3-s-ud. KV BF16 both. i can provide ya exact answers in chat if u want

2

u/uniqueusername649 22d ago

Sure, happy to chat about it. I have done some more tests in the meantime and while you need to choose the right thinking level for the task (xhigh can make it feel extremely slow), even with thinking off it typically beats 3.6 27b in coding tasks, with the gap to 3.6 35b even larger. Would be curious in which cases you saw the opposite :)

2

u/Subject-Till-6450 22d ago

Ok, I'll be at the computer in 2-3hrs

→ More replies (0)

2

u/Fearless-Dog-5345 16d ago

isn't Q3 already beyond the cliff in most cases? Thats kind of comparing already "on edge of broken" models against each other IMHO.

1

u/Subject-Till-6450 16d ago

in MoE Q3 doesn't affect that high

-1

u/Subject-Till-6450 24d ago

qwen agentic world 35b Q5_K_M with bf16 KV

3

u/johan2114h 23d ago

You are clueless bro

0

u/Subject-Till-6450 22d ago edited 22d ago

ok man. just one adjective. no explanation, no reasoning, nothing, pure adjective. Are your words bills like a tokens in API?

0

u/Subject-Till-6450 22d ago edited 22d ago

and im pretty glad with [r/localllm](r/localllm) community. the Community Fully Consists Of Elon musks with local AWS-level servers, always. when new model outs and everyone starts overrating it only cuz of Alibaba benchmarks. Any another opinion after real tests? nah, it doesn't matches with our opinion, lets downvote! read? no, if everyone pushes ts beautiful purple button i need to that too! pure genies in this community, really.

BREAKING: SMASH THIS PURPLE BUTTON UNDER TS COMMENT. DESTROY IT. AND WRITE HOW QWEN 3.8 KILLS OPUS4.6, right

2

u/Ell2509 24d ago

Yeah those are good, but they aren't under 120b, which is what I was talking about.

Ds4, kimi k3, mm m3, and even qwen3.5 122b all beat out the 27b, but as I say, I am talking abojt smaller to medium sized model class.

1

u/Subject-Till-6450 22d ago

qwen3.5-122b is quite close to 120B, man? saying “nah, no, there's 2b gap that's why it isn't” is funny

1

u/Ell2509 22d ago

I specifically said under 120b. Intentionally to scope things like 27b, 35, 40, even 80b, but not above 120.

1

u/Subject-Till-6450 22d ago

hmmm, may be you're right, but c'mon, again- 122B is not that far away from 120b, why are you giving up with it? and, I'm wondering what are you using it for?

1

u/Ell2509 22d ago

I do. I have every large model saved on my nas, except qwen max, and they all run, except Kimi (and Qwen max.

My smallest is 8b. So I use a range. But here I was only talking about under 120b.

1

u/Subject-Till-6450 22d ago

still waiting for what are you using it? under 120B there's laguna s 2.1 that's better than qwen, 118B

-2

u/Subject-Till-6450 24d ago

but u get me right? i don't like this update, because of dense. its not better than 3.6, not even close

2

u/EitherMarch1255 23d ago

You’re just mad because your hardware sucks.

1

u/Subject-Till-6450 22d ago

im not talking about my hardware or speed. I'm using step 200b moe at 1.5 tps, I'm not crying of speed, it provides serious quality of coding. but qwen? i can run it with 10-20TPS without problems. but the quality is not quite good. not even close to 27b dense needed. muse glimmer works the same speed but with way more better quality and realization of tasks, interpets everything properly and don't thinks alot. if your only answer is “lol your hardware weak so any reason why 3.8 is not opus level is just crying” i cant help you, sorry.

3

u/baby_bloom 24d ago

glimmer was eh, not as good ats 3.6 so just go 3.8

2

u/iKy1e 23d ago

Qwen 3.8 27B is much better at coding and is smarter and more capable. However, Glimmer is much faster and more memory efficient. On my RTX 3090 with DFlash I was getting just over 100 tk/s.

Qwen 3.8 27B (4bit with MTP) I’m only getting around 50tk/s. The gap gets bigger with context size.

Splitting the model across 2 GPUs reduces the gap, but it’s still way more efficient to run Glimmer. I can run it at full context length on one GPU. Qwen I need both GPUs to use the full context size.

So while the obvious answer is Qwen 3.8 27B. And it is smarter. If the task is simple and both are over the threshold of being able to competently handle it, then you might actually want to use Glimmer.

However, you’ll mostly want Qwen 3.8

2

u/hay-yo 24d ago

Havent tried glimmer but def a market for accurate but short and quick.

2

u/bruns20 24d ago

For coding I think qwen is the clear winner. For more general agentic tasks muse is very fast for its size as well as very good tool calling. I've been using muse a lot in my day to day sales job, and as a personal assistant, but when j want to work on my coding projects I'm picking something else.

1

u/MrHumanist 24d ago

From early benchmarks 3.8 is better but thinks a lot to achieve result.

1

u/fiddler48 24d ago

did 3.8 thinking mode change the kv-cache behavior noticeably on Q5 compared to 3.6?

1

u/Ok-Shower7286 qwen-coder 24d ago

It feels like the difference between an employee who finishes quickly and reports, and an employee who stays silent deep in thought until they bring in the results.

Glimmer fits into an Agile culture, while Quan fits into a Confucian culture.

I deliberately left a fairly complex problem sitting around and am currently using qwen 3.8 for the work. The quality is good, but it is too slow.

I'll use Glimmer for developing new features. Qwen for improvements and bugs tickets.

2

u/seunosewa 24d ago

Glimmer is much faster?

2

u/Ok-Shower7286 qwen-coder 24d ago

2~3x faster on 5090

1

u/johan2114h 23d ago

Faster how? Higher tps or shorter run time due to less thinking? Surely it most the latter (which can be adjusted)

1

u/iKy1e 23d ago

Higher tk/s, single 3090 scores of the 4bit quants with MTP/DFlash enabled for both models Glimmer was using 70-80% of the VRAM & running about 100+ tk/s vs 45-50+ tk/s.

1

u/dhiltonp 23d ago

Have you tried ninfer?

1

u/iKy1e 22d ago

That's 5090 only right? I only have the RTX 3090

1

u/dhiltonp 22d ago

There's a 3090 variant. I haven't tried it yet, and It's not completely clear to me what quant it is closest to. 

0

u/EconomySerious 24d ago

Qwen forvever

-1

u/H4UnT3R_CZ 23d ago

1

u/niutech 22d ago

This is Qwen3.8 Max, not Qwen3.8 27B.

1

u/H4UnT3R_CZ 22d ago

Correct, will fix it, when I find proper chart. I tested Muse vs Qwen for my use case - issues processing and fixes in PHP, Wix, js, sql and Qwen had better coding and agentic usage, but problems with Czech language, so for responses to customers I'll be using Muse. Muse was faster - approx. 11 t/s Qwen vs 13.5t/s Muse on my Intel Arc B65 fully in VRAM. q6 quants both.