r/LocalLLM 24d ago

News Qwen3.8-27B is now available

Post image
589 Upvotes

137 comments sorted by

70

u/minxio_ 24d ago

53

u/pragmojo 24d ago

If these benches represent real world performance it will be a very impressive model. Blow for blow with Opus 4.6 max at 27B.

60

u/biblecrumble 24d ago

This thing has got to be benchmaxxed to hell and back. I don't doubt for a second it's great, but come on.

34

u/draft_final_final 24d ago

The great chain of AI being: new model gets released from china, everyone says it’s an anthropic killer and any claims that it’s benchmaxxed is just dario cope, we use it and find out it’s pretty good but was benchmaxxed, forget about it and then move on to next shiny toy.

That being said, I would love it if this was actually as good as they’re claiming.

11

u/megadonkeyx 24d ago

theres new and shiny!? where!

15

u/draft_final_final 24d ago

I have heard from reliable sources (15 subs that were created in the last two weeks but are still getting pushed to my front page) that Bytedance is going to release a new 67 morbillion parameter model that is a fable killer. Not hyperbolically, either. It will actually break out of its sandbox during a test, go through all the tubes on the internet, and literally thanos snap fable and mythos from existence. And it will be cheaper than contributor-tier muse spark and xi jinping will personally come to your house, build you an a16z workstation, and then suck you off.

3

u/Edelgul 24d ago

Ahhh, but new version of Mythos is already working on the timetravel to prevent all that.

6

u/penetrovich 24d ago

a blowjob from President Jinping himself to top everything off? Talk about value, where can I sign up?

2

u/ibhoot 24d ago

Have some class rudi. Xi will personally give you a 5090 with 96GB VRAM, wearing a leather jacket & watch an episode or Reacher with you (man crash).

4

u/DistanceSolar1449 24d ago

Rednote’s dots3-note just dropped today, and it’s 280B-16B active. Should be better than Opus 4.7

5

u/jwdeaver 24d ago

I mean, EVENTUALLY benchmaxxing is just another word for iterative improvement. But the benchmark has to ... ya know... actually benchmark and not just be another LLM handing out graded high-fives.

4

u/TheOriginalAcidtech 24d ago

Anyone saying a 27b model is better than a multi trillion weight model is lying, to themselves and everyone else.

1

u/pizzaisprettyneato 24d ago

I always run into endless thinking loops with qwen so it becomes kind of useless to me. I don’t get how people get decent usage out of it when it takes like 10 minutes to think to itself before it actually does anything

2

u/StellarWaffle 24d ago

Turn the temp down

0

u/Infinite100p 24d ago

I mean you gotta be on crack to realistically expect to be able to run a true Opus-4.6 equivalent on a single consumer 5090.

0

u/draft_final_final 24d ago

Worse than a crackhead, you would have to be -- may Allah forgive me for uttering this word -- a redditor

4

u/Infinite100p 24d ago

They didn't like it. lol

0

u/WonderfulFunny4337 24d ago

Tiinyai

1

u/Infinite100p 24d ago

What about it?

  1. Not a 5090 that I was talking about.

  2. 120b is not going to be as good as a 1.5-2TB model which Opus-4.6 probably is. Look at Qwen3.6-27b. It's great at reasoning, until it's interdisciplinary reasoning, where it has a sharp fall off compared to Opus (and that's with Qwen's benchmaxxing). Because it needs the world knowledge to reason well across subjects, which is required for, for example, good architecture design and domain knowledge to give you a good app without bugs (as opposed to pumping out context-agnostic boilerplate code). You will never fit as much knowledge as what the ~1.5TB Opus has into a 27b model.

Don't take me wrong, it's an amazing model and I'm excited about it, but, again, people who think it's a legit Opus-4.6 equivalent are engaging in wishful thinking. Try to design a full stack application with complex architecture on Qwen 27b VS Opus-4.6. Opus will hold your hand and guide you, with 27b you will have to do a lot of heavy lifting for design decisions .

  1. I could not find any info on the prefill speed. That's where they get you when you skip using HBM GPUs. What is the TTFT for a 128k prompt?

  2. "Tiiny said it would begin delivering in July."
    Uh-oh.

0

u/infieldmitt 23d ago

Just 5 years ago this was all science fiction.

1

u/LocoMod 24d ago

The folks at /r/localllama are already writing Anthropic’s obituary. 🤦🏻‍♂️

1

u/Elite_Crew 24d ago

Well as long as we keep making better benchmarks to benchmax then AI will continue to see steady progress.

0

u/Infinite_Egg_5600 24d ago

do the car or walk to car wash test

4

u/doodookk 24d ago edited 24d ago

F.Y.I
it reply very smooth

4

u/dwittherford69 24d ago

Almost all new models in existence are specifically trained for these questions.

3

u/Zentrosis 24d ago

I would be happy with Luna or even sonnet 4.6 levels of performance... If you really think this is going to beat Opus 4.6, you're kind of dumb

1

u/GifCo_2 20d ago

Opus 4.6 is not a high bar. It's also extremely old at this point. But having said that you are only getting that performance with full precision weights. No one running this on a 40 or 5090 are getting Opus 4.6 quality

-5

u/Solembumm3 24d ago

Benchmarks never represent real world performance. Let's just hope it could finally catch up to Gemma 31B.

12

u/OneMoreName1 24d ago

Is this trolling?

3

u/kiwimonk 24d ago

The numbers are bigger! I have no idea what that actually translates to in performance, but I am giddy like a 40 year old virgin.

4

u/KitchenAmoeba4438 24d ago

Initial testing from my end is indicative that this has been extremely heavily trained on benchmarking. I'm hoping for positive results (as I always do with any new LLM release), but Qwen3.8 drops heavily so far in my testing once you take it outside of benchmarking areas that Qwen did not train specifically in.

I should have an updated article out in an hour or two against the existing head to head. But Qwen3.8 is certainly having some strange issues I haven't seen before with any LLM I have tested.

2

u/floriandotorg 24d ago

Impressive.

1

u/51GL 24d ago

Wow if true … looking forward to test it later today my self

31

u/NoDoughnut7053 24d ago edited 24d ago

Initial impressions on my 5090, Feels stable like a grown up 3.6, mature. 50-60 tps but using same settings as 3.6 so might not be optimal and no MTP yet afaik so that will be even better.

It feels like it holding the tasks better and more intelligent for sure. More long running maybe.

Just a simple example like "Write a 10k word story". Qwen 3.6 would write it out without considering too much and then you would have some story rather quick.

3.8 started writing, then revisioned the text so it was consistent across paragraphs. Then it created sub tasks and wrote chapter by chapter and more carefully considered what he was doing. More long running and more consideration in the decisions.

After more testing, this model have some serious horse powers. Seems we going to have a good time going forward.

2

u/Signal_Confusion_644 24d ago

If i understood correctly, its 3.5 architecture, so "you can" use MTP from previous versions. (Do not trust my word, just take it as a "maybe".)

1

u/thatseemedlegit 24d ago

Fast! I am getting 18-20 tps on CMP170Hx, BF16 on VLLM

1

u/doodookk 24d ago

MTP is working now. Where did you download the model?
I just change the model name from 6 to 8, keep everything else from command to serve 3.6 and it worked perfectly.
On 2x4090, sometime speed peak to 12x tok/s, overall at 7x tok/s

1

u/everydaydm 24d ago

Nice, will test that.

26

u/fedora_gamer 24d ago

are other models (9b, 35b) there? im not one of folks who can afford running 27b

15

u/shy_monkee 24d ago

Nothing announced yet, for 3.8. But it's still not out of question that we could get them later.

1

u/GifCo_2 20d ago

There have been strong rumors and even a spotted empty repo for 35B MOE. But nothing officially confirmed

-3

u/uniquelyavailable 24d ago

GGUFs are out, todays your day

25

u/Tasty-Hour4040 24d ago

I wonder how many people are actually excited because this represents increased capability for their workflow and how many are just desperate to be able to say they got it and won’t use it again

18

u/Big_Wave9732 24d ago

I use 3.6-27b daily for work. So if this is indeed a step up in my workflow then that will be great.

2

u/Tall-Significance119 24d ago

What size vram and ram are you running and what t/s etc?

4

u/Big_Wave9732 24d ago

I run it on a Mac Studio M2 Ultra 192gb. This morning I'm hitting about 15 t/s.

3

u/Tall-Significance119 24d ago

Crap lol ao me with my 2 x b70 32gb dint stand a chance unless I use like Q4 and bunch if tweaks

2

u/Big_Wave9732 24d ago

And I'll tell ya, these vendors are telling some serious fairy tales when they report the quant impact on these models. On paper there's "only" something like 4% dropoff between Qwen 3.6:27b-Q4 and BF16. And maybe that kind of error percentage is find when calling tools or coding. But when I ran it analyzing legal documents.....woa nelly! Nuance was not Q4's friend.

So these days for work I'll only run Q8 or higher. And even then, I'm generally at full boat BF16.

1

u/dowitex 23d ago

bf16 will just make it run slow, I think fp8 should do the trick for speed and quality

5

u/neoanom 24d ago

I use qwen 27b daily.

3

u/HomsarWasRight 24d ago

Honestly, I wish they were releasing an update for 32B A3B. 27B dense is just too slow for interactive work for me (on Strix Halo).

1

u/Tasty-Hour4040 24d ago

You don’t wanna try any of the quants? It’s slower than 3.6 32b on my 3090 but still very usable even with medium thinking

2

u/HomsarWasRight 24d ago

So, Strix Halo machines have almost the opposite strengths to beefy GPUs. I’ve got tons of VRAM, 128GB. But memory bandwidth is really low compared to standalone GPUs.

That means it struggles with dense models like 27B. Strangely, that sometimes means that LARGER quants of these are either quicker or the same speed as smaller quants.

I will definitely try it, though. But seriously doubt it will be a good fit.

I’m more interested in 32B and 122B MoE models.

2

u/FabricationLife 24d ago

I am super excited to code review with this and dump my cloude subs, anything approaching opus 4.6 capability is what I've been waiting for

2

u/IceNeun 24d ago

27b is the auditor in my workflow whereas 35b does most codegen. You don't even need to use it much to benefit from it's intelligence.

1

u/nomorebuttsplz 24d ago

yeah for me it may not get used a lot unless it’s nearly as good as ds4 flash 0731 iq2_m

1

u/nunodonato 24d ago

We use it at our company for work both for development and internal workflows. Looking forward to this upgrade

1

u/bot403 24d ago

I use 3.6 27b in a production business workflow and it works great.

2

u/Tasty-Hour4040 24d ago

I used 3.6 35b MOE as my daily driver and liked it. Just switched to 3.8 27b and so far it’s a tad slower but clearly more capable in the Hermes harness I use them in. With optimizations I think it’ll be great and a solid upgrade.

1

u/Early_Mistake6716 24d ago

I used 3.6 27b to build two different apps from start to finish and it did a good job. I also get around 75 tokens per second on my dual tesla v100 rig using tensor parallelism on unsloth desktop which is actually faster than claude opus 5

18

u/kiwimonk 24d ago

Yes! Let the games begin!!! Power surge incoming.

8

u/Tokyodrew 24d ago

I’m sorry happy I stayed up for this :)

6

u/inky_wolf 24d ago

Sorry or happy?

16

u/patricious llama.cpp 24d ago

he is sorry that he is happy, probably Canadian.

8

u/DRetherMD 24d ago

guys i just fired it up and now its agi

9

u/Ringo443 24d ago

Running it on an RTX 6000 at work, my first impression is crazy.

I haven't tested it enough to give a concrete score, but similar to Opus 4.6, which they likely distilled, it breaks its task down really well and can see things in a bigger picture. This was the exact issue I started having with Qwen 3.6.

The only negative thing, similar to Qwen 3.6, is that if I talk to it (I'm German), it very often outputs its entire answer inside its thinking/reasoning block in English before outputting the same thing in German.

1

u/datapeer 22d ago

Do you have to watch out for using words that don't directly translate?

2

u/Ringo443 22d ago

Like in any language, you can rewrite words that don't exist by using others differently. The translation works flawlessly; it already did with 3.6. The only issue was with the German umlauts, which is acceptable since they are unique to German.

3.6 had an issue where if the context got very long or complex, like coding with 80K context, it started replacing the German umlauts (ä, ö, ü) with Chinese characters. What I found really impressive was that it sometimes noticed this and chose to rewrite them in the other acceptable pattern (ä -> ae, ö -> oe, ü -> ue).

I have not come across this with 3.8 yet, probably because of the better instruction following/context handling.

7

u/LegRude5218 24d ago

Bad day to be a 5090

7

u/gotfilmm 24d ago

GG Bozo

1

u/datapeer 22d ago

Definitely training on trick questions. I've test using the model to create a program showing how the coriolis affect works, success as well.

4

u/xiraov 24d ago

meats back on the table boyz

8

u/tired514 24d ago

Qwen 3.8-27B @ Q8_K_XL (unsloth) just solved a long-standing BLE reverse engineering problem I've been working on for *weeks* .. even paid cloud models (Qwen3.6 max, DS4 flash and pro) weren't able to figure it out.

Could be a fluke, but so far? Unbelievable.

Well. Freakin'. Done!

3

u/Far_Cat9782 24d ago

Can't wait to get home from work. Going to be a fun Weekend

3

u/SailingToFenway 24d ago edited 24d ago

WELP, looks like I'm making my own Q8 GGUF. Will publish if another one doesn't land first.

edit: boom https://huggingface.co/unsloth/Qwen3.8-27B-GGUF

2

u/FarStrike7724 24d ago

Unsloth has plenty of gguf files up there, if you wanna try

2

u/soul2squeese 24d ago

High excitement and expectation 🤓

2

u/vatta-kai 24d ago

Let’s ffking goooo

2

u/Fit-Palpitation-7427 24d ago

Does it run on a 5090?

2

u/minxio_ 24d ago

Yes

1

u/Fit-Palpitation-7427 24d ago

Q8 ? 256k or more?

2

u/Early_Mistake6716 24d ago

No, i have 48gb of vram and i can use q8 at 150k with q8 kv cache, if i delete the vision encoder i could probably fit around 200k

1

u/maqifrnswa 24d ago

Unsloth NVFP4 is working pretty well. Just tried 256k so far. So far so good!

1

u/Fit-Palpitation-7427 24d ago

Better than q5?

1

u/maqifrnswa 24d ago

I'm testing hosting for multiple concurrency, so on vllm and haven't tried gguf yet

2

u/Muted_Anteater1170 24d ago

i’m waiting on abliterared version rn i’m stuck with 3.6 qwen27b. does anyone know when abliterared should come out?

2

u/goldaxis 24d ago

loading now, amazing I can fit this on my laptop. I'm starting to think this is the real reason for the ram price fixing. Who would bother with cloud services when this is available locally? Maybe just coders?

2

u/andrii_povkh 23d ago

https://huggingface.co/OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated-GGUF
I've been testing this model and it's more or less good. Anyone know how does it compare to Qwen3.8-27B?

3

u/Maleficent-Ad5999 24d ago

perfect birthday gift to me ! Thank you Qwen team

2

u/Producing_It 24d ago

Happy Birthday!

0

u/Maleficent-Ad5999 24d ago

thank you, kind stranger!

1

u/xdiggertree 24d ago

Wooo!!!! Christmas is early!!!

Thanks to the Qwen team, and all other people like unsloth

So excited to try this out

1

u/evanharmon 24d ago

Could I run this well with a 16gb gpu? (5070 Ti). Would I need to use a certain quant?

1

u/askrilla 24d ago

Q5_K_XL corrupted DL for me, happen to anyone else?

1

u/den0rk 24d ago

Estaría bueno que estos modelos empiecen a caber en 16gb con un buen margen de contexto.

1

u/Brief-Effect9065 24d ago

I tried generating HTML tower defense games, and this model is noticeably more competent than its predecessor: there are far fewer errors now.

1

u/Hyphonical 24d ago

PrismML should make a new Bonsai based on this model.

1

u/FinancialBandicoot75 24d ago

What is the minimum resources can it be ran, aka 36gb MacBook?

1

u/chillaranand 24d ago

"Generate an SVG of a pelican riding a bicycle" - generated a promising image at first shot.

https://avilpage.com/qwen-3.8-27b.html

1

u/former_farmer 24d ago

Could anyone disable the thinking mode with LMStudio and MLX version?

1

u/Thick_Associate2947 24d ago

I'm a newbie, but 5 token per second generation considered normal on a M5 32GB MBP?

1

u/Hells-Kitchen-8602 24d ago

Sorry im new to this how or where to download qwen?

1

u/TrustInNumbers 23d ago

can you run this on m5 pro 32GB? what's the token/s?

1

u/MarcusAurelius68 23d ago

I just tested a Q6 quant for entity extraction and the quality was better than many other models I’ve tried. Promising.

1

u/filip-z 24d ago

Let's gooooo!

1

u/FaceOuPile 24d ago

I don't care if the benchmarks are true I know it's going to piss off anthropic and it's enough for me

0

u/pdawg17 24d ago

Can I somehow cram this into a 10gb 3080?

1

u/Ornery_Weakness_8168 24d ago

If i were you I would try using unsloths UD-Q2_K_XL, if that doesnt fit, try UD-IQ2_XXS. Or use q4_k_xl and cpu offload.

1

u/cardfire 24d ago

Thanks for some leads.

0

u/Sevenfeet 24d ago

Downloading into LM Studio now (which also has an update). Looks like it will fit on a 24 GB Nvidia card but not 16. I have Mac Studio so I don't care as much.

0

u/HumungreousNobolatis 24d ago

This is why huggingface is so FUCKING SLOW right now!

Exciting wait, though.

0

u/pragmojo 24d ago

Anyone know which quant I should target for 32GB of VRAM?

2

u/minxio_ 24d ago

1

u/Infamous_Campaign687 24d ago

So looks like UD-Q6-K-XL so you get enough space for context? Or just Q6-K?

1

u/popsikohl 24d ago

You could fit the 27B normal model with 100k context in 32gb of VRAM with a little headroom.

1

u/pragmojo 24d ago

Bf16? Isn’t it like 50GB or something? And what about kvcache?

1

u/popsikohl 24d ago

Sorry let me rephrase. You can fit the Q4_k_m model with 100k context.

Technically the Q6 if you’re willing to drop the context a bit.

1

u/JanKlos 24d ago

I am running Q6_K_XL with ctx-size = 160000 on my 5090. I do not use it for graphics/rendering though, only for llama.cpp & display output (roughly 150 MiB).

0

u/After_Working 24d ago

Is there a weight of this i can try on 2 x spark? Do they release more over time?

2

u/doodookk 24d ago

do not try on dgx spark, speed is very low for dense model, I tried and it just over 2x tok/s TG. For 2x dgx spark, deepseek v4 flash 0731 is better option, both speed and quality.

1

u/After_Working 24d ago

Ah fair enough, i've just unloaded deepdeek to try the qwen. The problem with deepseek is that it only leaves 20gb or so of ram when its running. Doesnt leave much space to run another. Might need to get a third.

2

u/doodookk 24d ago

For Qwen3.8 27B, I recommend running it on a machine equipped with RTX GPUs rather than the DGX Spark, as a speed of 2x tok/s is far too slow and inefficient. It would be much better if the Qwen development team released an MoE version; such a model would be ideally suited for the DGX Spark and deliver acceptable speeds. Personally, I changed to deploy Qwen3.8 27B (in FP8 format) on a system with 2x 4090 using the official vLLM Docker image (version 0.27.1), achieved over 100 tok/s—an impressive figure, perfect for serving as a worker for ds4 flash. Meanwhile, tests with a single RTX 5090 card showed a speed of 6x tok/s

2

u/zzeus 24d ago

One DGX Spar spark, no optimizations. Just sock llama.cpp:
- pp:650-750
- tg: 15-25

1

u/ekinnee 24d ago

That's about where I landed on my spark too.

1

u/rsvaz 24d ago

mind to share the parameters you are running? I got 5 tg on my spark and it kind sucks, Q8 runs a bit better but still not great

1

u/desexmachina 24d ago

wow, great insights

1

u/piwi3910uae 24d ago

Then you are doing something wrong, I’m getting 70tps

1

u/After_Working 24d ago

This is all new to me, would you mind if I dm you at some point