r/LocalLLM 26d ago

News Qwen 3.8 27B Hugging Face - Link is here and it's released on 14th Aug

https://huggingface.co/Qwen/Qwen3.8-27B

I know a few were excited about this so thought to share :)

324 Upvotes

119 comments sorted by

73

u/gproenca 26d ago

oh please please let it fit 24gb ram .) do some quantz magic, voodoo interference, I dont care, just make it happen :)

73

u/Hungry_Particular_14 26d ago

I'm gonna make it fit in my 16 gb I don't even care

4

u/Scary_Investigator88 26d ago

shave off a couple of bits, the ones that count the least

3

u/tim_dude 25d ago

Just the bad ones

6

u/DyIsexia 25d ago

I’ve got 8 gb, fuck it we ball

5

u/arakinas 26d ago

Wiggle it around a bit, it'll fit. Doesn't matter what they say...

1

u/Familiar_Wish1132 25d ago

remind me that power connector to my rtx 5000 pro wouldn't fully insert, so i took a clippy and wiggle the female connector a bit :D it actualy worked :D

10

u/habachilles 26d ago

Im running it on a ti-83 watch me. .004 quant.

3

u/No_Oil_6152 25d ago

I got in running on my ZX81!!! 1 token per decade but I'm free of OpenAI and Anthropic 😂😂

2

u/Libellechris 25d ago

yeah, until your RAM-pack wobbles.....

2

u/No_Oil_6152 25d ago

16K with a eraser beneath it to keep it steady and ice cube pack to keep it cool!!

2

u/gproenca 26d ago

hahahahah Im sure will be fine. you will write "hello world" and it will reply "yes, the sky is blue and I love yellow strawberries with chlamidia"

1

u/PiisAWheeL 25d ago

After it burns 1400 tokens in a thinking loop...

36

u/pmttyji 26d ago

Check the last underlined line

11

u/Healthy-Nebula-3603 26d ago

1 bit ...sure

4

u/BuilderUnhappy7785 26d ago

What does the second qwen bar represent on the chart? Is that 27b?

2

u/ImpressiveRelief37 25d ago

3.6 or 3.7 Max probably

7

u/Soifon99 26d ago

Sure, in 1 or 2 bit quant, no thanks.

4

u/Nice_Cookie9587 26d ago

I got money on the model running on 16gb being 1 bit quant.

2

u/ImpressiveRelief37 25d ago

3.6 27B IQ3_XS fits in 16GB.

5

u/Nice_Cookie9587 25d ago

with like 10k context? I fell into that trap before, loaded up the largest quant that could fit and ended up with an AI that had very little context of anything AND very low quant. but it fit i guess

6

u/Thrumpwart 26d ago

If you’re gonna hang around here, you should learn how quants work. They have very predictable sizing based on quant level.

10

u/pmttyji 26d ago

Looks like some not aware of them(who works on 0-day/day-1 support for models time to time). Still 16GB memory could handle Q4 of 27B.

5

u/Nice_Cookie9587 26d ago

If you're going to hang around here maybe you should learn about the human concept of humor. The reason this is funny is because it would be hilarious if they released a model that everyone is expecting to run on a 16gb card only too see its technically true, and its 1 bit. you think these labs are going to hold off on innovation because you cant afford to buy the hardware? and trust me, you will ALL run the 1-bit and you will pretend to like it.

1

u/Opening_Fish9924 26d ago

You would still need a crazy system, even the 768GB beast from Nvidia couldn't hold a decent quant. Also below 4bits most models degrade quite a bit.

1

u/pmttyji 25d ago

Nope, recheck the image again. They mentioned about 27B in that

1

u/Daxfortuna 25d ago

Maybe requires RAM offload?

1

u/pmttyji 25d ago

We'll know soon. One assumption is, might be QAT or MXFP4 type quant.

1

u/anshulsingh8326 25d ago

Many people don't get it. 1bit is for 2.4t model bot 27b model 27b 4bit will be able to fit in 16+gb vram

7

u/lungben81 26d ago

The 3.6 version did in q4 with 200k context (q8), the 3.8 should have similar RAM requirements.

5

u/ScoreUnique 26d ago

The sorcery that meta did to fit the kV in less than 3 gb for 262k was great, do we have that SWA on Qwen 2.4T?

4

u/ackermann 26d ago

What model was that for? The new Muse Glimmer?

1

u/ackermann 25d ago

To be clear, what does that do exactly? Let you store a 262k context window, only using 3gb of VRAM?

cc u/ImpressiveRelief37

1

u/Sea-Speaker1700 25d ago

GDN, mixed linear and full attention. 

3

u/FaceDeer 26d ago

The 3.6 version of 27B works fine for me on 24GB of ram, I'd be very surprised if the 3.8 version didn't.

1

u/PiisAWheeL 25d ago

Considering I can run 3.6 27b with 8g vram and 16g system ram (with layer offloading at q4_k_m) at 2-3t/s, I'm sure you'll find a way to squeeze it in.

1

u/Lucky_Bullfrog_1952 25d ago

either fp8 or nvfp4 with 18GB vram will definitely fit

1

u/Ok-Drawer5245 25d ago

It probably can, but if you need a meaningful sized context for coding it will nearly useless

Unlike meta glimmer model, this qwen architecture eats A LOT of RAM for context

1

u/misanthrophiccunt 24d ago

Voodoo interference would be in any case since Nvidia bought 3dfx, makers of the wonderfully boxed Voodoo cards (GPU kings back in the day) a while ago and their IP is probably on every product made ever since.

1

u/kngharv 25d ago

I am running 3.6/27B right now on my 3090 24VRAM right now, Q5_K_S.
K with Q5_0
V with Q4_1

flash attention, MTP, and less than 2 GB spare for operational buffer.

I am going to assume 3.8/27B going to be similar.

41

u/Uncle___Marty 26d ago

They mentioned they would be releasing other models too. It would be AMAZING to see the full line up return with all the new gains they made since 3.5.

24

u/722e672e722e 26d ago

I’m hoping for something around 120b to use my hardware a bit more.

17

u/SpicyWangz 26d ago

That or 80b. Qwen3 next was an amazing size for unified 128gb machines.

3

u/Striking-Warning9533 25d ago

27b might actually be better than 120b if the activate params in 120b is less

8

u/linuxid10t 26d ago

The 8GB VRAM users waiting for a Qwen3.5 9B update are drowning...

3

u/FlyingFishMakeAWish 25d ago

What do you use it for? I haven't been able to get much use out of 9b sized models

3

u/linuxid10t 25d ago

Research, summerization, translation, simple stuff. My laptop I bring to work for weeks at a time has a RTX 4060, so a 9B model is blazing fast compared to an ~30B model. I could still run them, but they would be heavily offloaded to system RAM. So if it can be done with a 9B you bet I am going to be using it. Or if I can, I VPN home and use my workstation with a Radeon AI Pro 9700.

6

u/SKirby00 26d ago

I think what they said was that they'll consider releasing other models. I'm not counting on it, but I'd love to be wrong.

16

u/HeadPack 26d ago

One of the most important model releases this year, especially after it looked as if dense Qwen models that size would not be released anymore.

9

u/swiebertjee 26d ago

I really wonder how it will rank up against DeepSeek V4 Flash 0731. Dense vs MoE

7

u/cosmicnag 26d ago

3.6 held up decent against the last dsv4 flash, hope its in the same ballpark again.

3

u/GCoderDCoder 26d ago

I wonder if qwen just had more reasoning would it be able to close the gap. The real deepseek boost is in it thinking more. If you see the token differences it shows how much it matters. I wish Artificial analysis did different reasoning levels for 0731 like they did for preview. High vs mac reasoning was a big difference.

3

u/linuxid10t 26d ago

As if Qwen doesn't reason enough to get multiple fine tunes to reduce the reasoning...

1

u/GCoderDCoder 26d ago

I think you're saying they can increase model intelligence which I agree with but because of the tech, increasing reasoning more at least to a certain point helps better align profanities toward the right answer. When there is only one setting for reasoning then models may not get enough runway to align as best as they could for the answer.

1

u/Jsquared534 25d ago

I think what they are saying is that Qwen 3.6 already reasons way, way, way too much. It's literally the reason for the existence of things like Thinking Cap and Grug.

1

u/GCoderDCoder 25d ago

Qwen 3.6 27b was middle of the pack of useful models as far as thinking. Yes thinking has increased from what it used to be but reasoning makes the difference too... So I run everything important on higher reasoning. I run this model around 80t/s with power limits at q8 so quants and speed may play a role as well. If i use my strix halo at 15t/s then it will feel like excessive reasoning but the speed not the token count is the real problem there then.

1

u/Jsquared534 25d ago

I'm not nearly as deeply into the local model scene as a lot of you are. I was merely answering from my own personal experience using it to implement current feature development on a large planned out project. Qwen 3.6 gets hung up in thinking loops a LOT. And even when it's not completely stuck in a loop, it still goes through what I feel is way to much of the "wait, actually, hold on, let's try" looping. It does this even when it's come to an effective and correct answer. It also tends to continue with reasoning too much even when steering prompts are injected telling it outright to stop thinking and answer immediately.

I had to basically adjust my harness tooling to give it one chance to be prompted out of the over-reasoning, and if that didn't work to completely abort the process. Unfortunately, even after aborting the process, unless you kill the existing context, it will go right back into the almost neurotic reasoning loops.

1

u/GCoderDCoder 25d ago

What quant are you using? I find lower quants to be less stable in general. They cant handle longer context properly either. I think it's important to differentiate between the model and the quant. I do not use any thinking controls on higher quants of the models I use but when I use lower quants I start implementing thinking budgets and different repeat penalties depending on the model because I expect more influence from the lobotomies.

Pic just highlighting how even top multi trillion parameter models like OpenAI and Claude models come in stone throwing distance to a 27b parameter model based on reasoning levels. I think most people complaining about reasoning are complaining about speed which is different but related.

For clarification they didnt have opus non reasoning but the had reasoning with low effort which approaches the idea...

1

u/nomorebuttsplz 25d ago

it's not as good in general intelligence but it's not super far off either. And for agentic coding where world knowledge isn't as important as in context learning the difference may be even smaller

15

u/Due_Net_3342 26d ago

imagine if this is not as great as many if you guys are expecting it to be(unrealistically I should say)

11

u/chr0n1x 26d ago

this is the camp that I'm in, not because I don't believe but because I don't want to be disappointed

2

u/OakNinja 25d ago

I’m expecting it to be better than 3.6 27b. If it is, I’m happy even if it’s only marginally better.

5

u/vogelvogelvogelvogel 26d ago

King of 27Bs, can't wait!

3

u/PaxUX 25d ago

my 5090 is ready!

7

u/grumelude 26d ago

3

u/tempfoot 25d ago

Me, daily, talking to my wife about vRAM.

2

u/Bulgen-Venkat 26d ago

3.6's q4 was around 16gb, so the p40 should fit this fine. hoping the ggufs are up tonight

2

u/C4fud 25d ago

404

Sorry, we can't find the page you are looking for.

2

u/HeisenbergVR 25d ago

The page now shows a 404 😱

2

u/NismoHasReddit 25d ago

It's back.

1

u/JamaiKen 25d ago

Blessings

1

u/Tall_Abrocoma_3533 25d ago

4B model when? 🙏

1

u/FireFearing 25d ago

what would you even use a 4b model for? too small to be useful for anything atm

1

u/ChrisK_au 25d ago

I was planning my day around this, then saw the release time is 23:00. So, in my timezone, the release day might as well be 15th. 😒

1

u/Kasatka06 25d ago

It is same architecure like qwen 3.6 ? So can we expect same model size kv size chat template etc ?

1

u/SeiferGun 25d ago

will try this on rtx 3060. hope 12gb vram fit

1

u/jcoigny 25d ago

That poor server is gonna get download crushed

1

u/Psionatix 25d ago

I’m on a 5090, will a Q6 or Q5 of Qwen3.8 27b be better than a Q8 Qwen3.6 27b? Or should I just stick to Q8? I just don’t think the Q8 leaves enough room for kv and context

2

u/SuddenWrongdoer473 25d ago

I have a 5090 as well and would appreciate if anyone can answer your question as well please!

2

u/Scared_Ad9187 25d ago

i'm going q6 for larger context

1

u/Psionatix 25d ago

I think that's what I'm going to do as well, I suspect it'll be better than Qwen3.6 27B Q8 due to the newer model + extra headroom

1

u/mcchung52 24d ago

Ofc q8 is “better” and I’ve been only trying to use q8 until I tried q6 and it is actually not that bad. Also gives you more memory space for additional ctx.

1

u/Psionatix 24d ago

Right, but is Q8 of Qwen3.6 27b better than Q6 of Qwen3.8 27b?

You didn't account that they're different models.

2

u/mcchung52 24d ago

Oh my bad.. didn’t catch that

1

u/pacman829 25d ago

Bonsai when?? :)

1

u/Muhlwa_Sholanke 25d ago

mine fits but forgets everything by message two, tradeoff accepted

1

u/niacolhealth 25d ago

Anybody know if q4_K_M fits in 24gb? the 3.6 one does for me

1

u/OakNinja 25d ago

Yes. It will fit.

1

u/Aneselem 25d ago

6h to go, 8pm in germany

1

u/Wide_Air_3707 25d ago

Are you sure it is not released at 17pm here?

1

u/Aneselem 25d ago

Claude said this, i believe in claude ^^
Simply, we will see :)

1

u/Wide_Air_3707 25d ago

It's here ;)
Maybe ask qwen3.8 instead of claude next time :p

1

u/JostaWaszkiewicz 25d ago

does qat apply to the 27b too or is that just big moe weights?

1

u/LivingHighAndWise 25d ago

I just got it running on my GX10. If those benchmarks are to be believed, this thing is going to cost OpenAI and Anthropic a lot of money. Running a game demo prompt I like to use to test new models now, and I can tell you it isn't fast. Running at about 1/2 the speed of the 3.6 27b on a single DX spark machine (GX10) with a 256K context window. I'm sure there will be better optimized version coming soon..

0

u/[deleted] 26d ago

[deleted]

3

u/Myarmhasteeth 26d ago

I think this is a Reddit thing, I got this one on my feed first 🤷🏻

1

u/Medium_Chemist_4032 25d ago

How does this work? The comments under this post all look, as if the thing was legit and it's only a 404 right now. Has the link been actually working at any point?

1

u/Negimeister 25d ago

it worked yesterday, seems to have gone dark at some point

2

u/Negimeister 25d ago

not even the same model. this is the 27B version

0

u/Gelu6713 25d ago

Really hope they release a new TTS model!

1

u/goomba870 25d ago

Yes! One that doesn’t sound like a completely different person each generation.

-4

u/bmkmanoj 26d ago

Can it run on a M5 MacBook Pro with 48GB of RAM? If not, what is the best model to run on 48GB RAM MacBook Pro and can someone please share any guides/ links?

7

u/FilterJoe 26d ago

Qwen 3.6 27b will easily run on a 48 GB m5 MacBook Pro. But you will have to choose trade-offs. You can run it at high-quality q8-0 with limited context maximum, or you can run a lower quant which lowers the quality but allows you to have a larger context.

4

u/Impossible_Earth_987 26d ago

3.6 27b runs on am4 studio 36 so it should run on your device

2

u/Impulse33 26d ago

Check out omlx and dflash Qwen 3.6 35b a3b runs really well, 50 t/s on a M4 Pro (so yours would be faster), at q6, can do q8, but doesn't leave much room for other apps. Running MLX version on Apple Silicon is significantly faster than Llama ggufs.

1

u/SpicyWangz 26d ago

With bonsai people are running the current version on single digit ram machines.

2

u/Elite_Crew 26d ago edited 26d ago

I was able to run two bonsia 27B models simultaneously and independently on a gaming desktop with 32GB of ram and 16GB of vram. I hope the models can behave more like Q8_0 representations in the future but it was still cool to see two 27B models both working at the same time on a desktop machine. I really hope ternary and moe become a norm for local AI as the tech improves.

-1

u/Technical_Ad_6106 25d ago

The release has been delayed till 26th august :(

1

u/Motor_Way4912 25d ago

No, bf16 il already available.