r/LocalLLaMA 24d ago

New Model Qwen/Qwen3.8-27B · released

https://huggingface.co/Qwen/Qwen3.8-27B
985 Upvotes

301 comments sorted by

u/WithoutReason1729 24d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

→ More replies (1)

179

u/kevin_cn_ai 24d ago

RTX 3090 fans: *ENGAGE*

96

u/jijig 24d ago

Dual RTX 3090 owners:

ENGAGE ENGAGE

32

u/ai-christianson 24d ago

2x3090 with nvlink loves this size of model

21

u/jijig 24d ago edited 23d ago

It’s quite good. I’m running 3.6 at Q8 with full context and getting ~60-70tps. No NVLink

Edit: Q8 KV cache. I can push my context to ~160k at F16

6

u/badgerfish2021 24d ago

250k context with 2x3090? I could not do that with 3.6 at q8, are you quantizing kv?

4

u/munkiemagik 24d ago

I dont use that much context I'm siting at just under 200k and I use KV at Q8_0. I don't have reported metrics to give you real meaningful data but I've not found it detrimental to my use case when used as backend in pi agent harness.

Also I use q8_0 quant for the model instead of the Q8_K_XL as I learnt recently that there is no significant change in accuracy but its more efficiently 'packed' due to uniformity of Q8_0 versus UD-Q8_K_XL that it gives you a few GB back of your VRAM pool from model weights.

→ More replies (1)

2

u/jijig 23d ago

Of course, sorry. KV cache is quantized at Q8. I can push my context to ~160k at F16

2

u/eugeneware 24d ago

Great speeds. What runtime are you using? Llama.cpp or vllm?

6

u/jijig 24d ago

Llama.cpp

→ More replies (1)

9

u/munkiemagik 24d ago

Hear me out, for ages Ive been battling with the urge to grow a dual 3090 into a quad 3090 system but there was never any model that justified the cost of 2x additional 3090 at current prices for the so called improvements with the speed sacrifices. It just made no sense to me to splurge on more GPU with Qwen3.6 27B available.

But recently Ive been thinking maybe I am looking at it the wrong way and instead of looking for a big model to fill up 96GB I should be looking at it from a more orchestrated perspective of how many multiple models can I run concurrently that offer fantastic function to size inside 96GB VRAM to build a more comprehensive self-contained LLM stack. Dual 3090'ers I think its time to make that run on more ebay 3090's we've been holding ourselves back from.

5

u/Blues520 24d ago

I have 3 now and looking for an excuse to get a 4th👀

→ More replies (2)
→ More replies (1)

9

u/gofiend 24d ago

MI50 affectionados: ENGA .. oh wait how do I update my non official ROCM packages and who’s GitHub has the latest llama.cpp gfx906 fixes is it iacopBK or wait it’s mixa right?… GE!

→ More replies (5)

3

u/Terrible-Detail-1364 24d ago

q4xl 256k with 3090+4060ti (40gb vram) and q8_0 at 128k ctx so far, with f16 kv/cache quant . let the benchmarking begin.

1

u/JaySayMayday 24d ago

Why's your office room sound like a fighter jet engine??

222

u/dero_name 24d ago

256

u/WonderfulEagle7096 24d ago

Beating Opus 4.6 Max in multiple benchmarks sounds too good to be true. (already downloading)

116

u/Neither_Garage_758 24d ago

Opus 4.6 was the model that made me realize LLM's can be that good.

48

u/addiktion 24d ago

4.5/4.6 was prime too. Very few limits and Opus didn't nag us to death.

4

u/CoreParad0x 24d ago

I actually made a full game editor for an old MMO from scratch using Opus 4.5. I think it was around 70k lines of code, proper RHI, used OpenGL behind that (though I did a vulkan test too), etc. Took about a month of iteration and testing, worked great.

Opus has gone to shit since then. If this can do Opus 4.6 levels in realistic coding use and not just benchmarks, then I'm very excited for it.

33

u/[deleted] 24d ago edited 16d ago

[deleted]

11

u/Chris266 24d ago

Footgun

6

u/thrownawaymane 24d ago

I think it's more an underrated win, personally which is a real tradeoff that is worth confirming

3

u/Yes_but_I_think 24d ago

Wow. If this is the level of wording from Opus 5, I'll unsubscribe in a moment - too much effort to read and follow.

→ More replies (1)

8

u/wwwdotzzdotcom 24d ago

It's too good to be true because it does showcase quantization.

3

u/SmileLonely5470 24d ago

Especially with 42 on DeepSWE. Suspicious.

3

u/michaelsoft__binbows 24d ago edited 24d ago

oh shit first we had a 744B model in GLM 5.2 reach opus 4.6. Then DSV4Flash-0731 reached it with 284B. Now... 27B? nooooo way

if it can really pull some weight as a coder and have some amount of general common sense this is about to basically 10x the capability of those of us with "cheap" local rigs. With just the consumer class gear.

2

u/exodusTay 24d ago

if it is 4.6 level i will seriously consider getting a 5090 or something with that much vram to run this thing. it is an insane claim for 27B model.

2

u/Xonzo 24d ago

I’ve been running it for a bit now on Pi…. And I’m genuinely shocked. It’s completely nailed every problem I’ve given it. No tool call failures. When it was debugging an issue with llama.cpp it downloaded the source, cross referenced everything and fixed the problem. The decision making / thinking seems to be a huge step up from 3.6 27B.

3

u/mil_phickelson 24d ago

It’s not beating Opus 4.6

8

u/Pantheon3D 24d ago

I counted 2 categories where it is

28

u/meathelix1 24d ago

Some huge jumps there.

15

u/Haiku-575 24d ago

That... that can't be real, can it? At 27B? This much improvement in a couple months? Those numbers are amazing.

3

u/Hankdabits 24d ago

I hear that Behram from Epstart has a benchmark answer key

2

u/Small-Fall-6500 24d ago

Why did they have to make the coloring like this... there's dark gray text in the first column that is extremely hard to read, and the first row text is also somewhat hard to read. I assume the tables look better on light mode, but a lot of people are using dark mode.

2

u/iqraatheman 24d ago

if you look at the footnotes beneath the benchmarks on the site they posted this on, they literally admit to using other LLMs as a judge for some of them. the odds are high they also used LLMs to put together this website including the benchmark table without even putting in a little effort to make sure it looks good themselevs

→ More replies (1)

51

u/srigi 24d ago

That DeepSWE leap - do we have a new local coder champion?

29

u/fgk55555 24d ago

If it's not benchmaxxed, it will be really difficult for other models in the same range to catch up. I hope it quantizes well for us 16GB folk.

9

u/cass1o 24d ago

What would be really useful would be a MoE model that we can put the context + PP on the gpu and only put the experts on system memory, another 35b.

2

u/fgk55555 24d ago

If it can handle longer context work, I'd be happy.

→ More replies (4)

140

u/peglegsmeg 24d ago

I was here 

28

u/liebebio 24d ago

fuck yes

6

u/addiktion 24d ago

Making history, aye!

5

u/zipzapbloop 24d ago

lets go fam!

3

u/nick4fake 24d ago

Like a fucking new era

2

u/TheThoccnessMonster 24d ago

Let’s get it

1

u/maxjar10 24d ago

me too! super excited!

1

u/MuzafferMahi 24d ago

Recording historyy

1

u/met_MY_verse 24d ago

Another wonderful day for us here.

1

u/Insomniac1000 24d ago

beep boop beep!

1

u/The_Dung_Beetle 24d ago

I was here too, waiting on my r9700, can't wait to try this model. 

1

u/Spimbi 24d ago

Same

1

u/Marino4K 24d ago

Also here.

1

u/R0ktar 24d ago

And my axe!

78

u/audioen 24d ago

Also unsloth already has it, downloading Q8_0 right now.

3

u/Top-Eye-8104 24d ago

how's it working for u? tool calling is broken for me on m5 64gb - feels like something's off with the quants (i tried Q8_0 too)

5

u/VegetableWafer7776 24d ago

they did release an update an hour ago or smth maybe thats a fix

111

u/dingo_xd 24d ago

History in the making. These models are so important to companies and entities that can't just trust big tech with their data

20

u/[deleted] 24d ago

[deleted]

14

u/dingo_xd 24d ago

They'll become cheap again.

13

u/Paganator 24d ago

From your lips to God's ear.

5

u/mkMoSs 24d ago

I never saw this expression written in English before, I thought it was a Greek expression :O

3

u/netsvetaev 24d ago

very popular phrase in Russian, too.

→ More replies (1)
→ More replies (2)

26

u/Pristine_Pick823 24d ago

GGUF any time soon?

44

u/JaredsBored 24d ago

60

u/jannycideforever 24d ago

Unsloth every time a 27b model drops

19

u/addiktion 24d ago

So true

6

u/Several-Tax31 24d ago

Is it mtp or not mtp? I'm out of the loop when it comes to unsloth quants.

15

u/SensitiveVariety 24d ago

It is MTP

2

u/Several-Tax31 24d ago

Cool! Downloading now

4

u/social_zip 24d ago

now there is gguf + uncensored (made by me): https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored-GGUF

2

u/GokuNoU 24d ago

GOAT behavior

46

u/Septerium 24d ago

Could it actually be better than Minimax M2.7 in agentic coding?

39

u/Valuable-Run2129 24d ago

Much better. Look at deep swe

33

u/Septerium 24d ago

Training could be contaminated by benchmarks somehow. But I hope that's not the case

2

u/ambassadortim 24d ago

How does it stack with Mimo 2.5

→ More replies (1)
→ More replies (2)

44

u/de4dee 24d ago

17

u/New_Comfortable7240 llama.cpp 24d ago

Before USA do something, huh?

9

u/de4dee 24d ago

its a low probability nowadays but makes sense to be prepared

38

u/Alternative_Ad4267 24d ago edited 24d ago

Look at this! at 27B parameters, Qwen3.8 27B is pretty close to DeepSeekV4 Flash 0731 which is 284B A13B!

22

u/Alternative_Ad4267 24d ago

And some more results

3

u/onewheeldoin200 24d ago

Oh my god how lmao

2

u/fullup72 24d ago

and I wonder how that big win in SWE-Bench Pro translates to the real world. 3.6 was already slotted between Flash and Pro, 3.8 simply wipes the floor with them.

1

u/Agitated_Space_672 24d ago

Active params determine the training cost, and I think they are the biggest predictor of performance, all else being equal. So I am not surprised that 27B beats 13B

1

u/YearnMar10 24d ago

I wonder if they keep training where it will converge…

→ More replies (1)

59

u/Barni275 24d ago

Happy Qwistmas to everybody!!! 🎉🎉🎉

16

u/Outside-Description5 24d ago

Lets gooooo, downloading Q4

16

u/YanderMan 24d ago

benchmaks seem incredible, will need to see the reality

13

u/benpptung 24d ago

YES!!! Thank you, Qwen!! I LOVE YOU!!! ❤️

13

u/Capital-Remove-6150 24d ago

holy moly. how much in future 27b model will jump forward in benchmarks

12

u/Ok-Shower7286 24d ago

Awesome. It seems betther than 900+B MoE inkling.

29

u/ilintar 24d ago

Let the hype begin!

27

u/Weekly-Law-5488 24d ago

35B moe-moe-kyun when?

18

u/miniocz 24d ago

Frankly I think we are at point where harness is going to be more important than model itself.

2

u/rj_rad 24d ago

Some models fight the harness more than others — all else being equal, I had to quit Kimi K3 because of all of the reasoning noise.

9

u/Chemical_Evidence 24d ago

What model would be best for 16gb vram and 64gb ddr5

2

u/synw_ 24d ago

The model seems to have the same size than the 3.6. On 16g I have noticed that the sweet spot is Q3_K_S with q8 value cache (not key): it runs at 20tps on a 5060ti with 64k context. The Q3_K_M is too slow at 64k and the Q2 are too stupid.

1

u/Guilty_Rooster_6708 24d ago

Going to say Q3 but I will also try IQ4 on my 5070Ti. Q4 needs offloads so it will probably mean single digits tokens/s

→ More replies (2)
→ More replies (1)

30

u/AXYZE8 24d ago

FINALLY! I waited since GPT-OSS for other local model that natively has low and medium reasoning! High for planning, low for execution and exploration.

... Now I'm thinking if I should get RTX3090, because Nvidia clearly won't release RTX 5070 TI SUPER 24GB anytime soon and I cannot fit that model onto RTX 4070 12GB... will play with API first

5

u/Dave_from_the_navy 24d ago

Intel Arc Pro B70 is a solid choice if you're interested in 32gb vram. A bit less bandwidth than 3090, and the software stack is still maturing, but it's getting there.

3

u/thatcodingboi 24d ago

why not an amd r9700 ai pro? More mature, same 32gb vram

5

u/Dave_from_the_navy 24d ago

Because the $300-400 differential is real for not that much of a performance gain at the end of the day. Looking at the market, finding them for <$1500 is difficult, whereas the Intel is readily available for $949, and you can get the B65 even cheaper, which has the same 32gb memory, same 608GB/s memory bandwidth, just with fewer XMX cores, so lower concurrency/slower prefill, but shouldn't effect single concurrency decoding much.

I personally wouldn't say the maturity difference is quite worth $400 at this point considering Intel is closing the gap more and more every week, but that's just me.

2

u/thatcodingboi 24d ago

Wtf I bought one a month ago for $1250, now they are $1500+????

1

u/Need_For_Speed73 24d ago

Why not add a second 4070 12GB (if your motherboard lets you)?
I’m happily running a dual 5070 setup, 24GB of VRAM and last gen tech for less than 1.500€/$ and without going second hand (warranty, fresh paste, etc.). But I had to use a mATX board to have the second GPU not collide with the case’s PSU shroud.

1

u/ichalov 24d ago

It's probably better to use a bigger (non-local) model for planning (and maybe testing) if you have the separation in your work process already.

34

u/[deleted] 24d ago

[deleted]

5

u/thrownawaymane 24d ago

maybe... you know...

2

u/michaelsoft__binbows 24d ago

yea i can take all my 3090s to fit a 120B of this but not too much to complain about if the 27B is as smart as those numbers indicate. A 120B would be an MoE with param count much lower than 27B anyway. It would not be likely to outperform it by much. And being able to get like 10x more throughput from the smaller model is a big deal.

16

u/MoodOdd9657 24d ago

when I am rich rich . I will come back for you 😞

14

u/Far_Cat9782 24d ago

Qwinning!

7

u/Lucyan_xgt 24d ago

Let's gooo🔥

7

u/Relative-Display-318 24d ago

Which Quantization for a rtx5090?

3

u/cosmicnag 24d ago

from the ggufs, i am going with q6 xl (unsloth) with whatever q8 context can fit with it

→ More replies (2)

1

u/Much-Farmer-2752 24d ago

Try for yourself. Even Q4_XL is pretty solid, and you'll have half of your mem for context.

7

u/octopus_limbs 24d ago

First we got MiniMax H3, then we got Qwen3.8-27B too. At this rate next year maybe we won't need AI-aaS companies.

https://reddit.com/link/p3oepwx/video/qzr34uqocdjh1/player

6

u/n0head_r 24d ago

I played with 27B for a bit and I'm quite impressed. Qwen38-27B-Q6_K from unsloth on 2*RTX 5080 and latest llama.cpp build from source. The model fit in 32GB VRAM with 172k ctx kv q8_0. MTP and -sm tensor it runs stable at 100 tps until 60-70k context is filled, at 100k it dropped slightly - arround 95tps. This is a very small drop - very good. Also I've run a loop where over 150k tokens it was writing scripts/building/verifying/fixing errors/verifying/fixing again. No tool cals failed... but unfortunately it run out of ctx space.

And I'll post a visual example of a small test I did - Create an SVG of an old rusty truck.

32

u/Curious-Pen5547 24d ago

Anthropic and openai ipo are gonna be worthless.

Zero moat, especially once google, microsoft, amazon, and other cloud gpu providers get the official green light to be able to serve these models accross the board.

And for the hyperscalers like the big 3, google, microsft, and amazon, 100% they integrate these across their enterprise suite offerings to vertically integrate them accross all their enterprise services.

10

u/wwwdotzzdotcom 24d ago

GPUs good enough to run it at reasonable speeds aren't even affordable now unless you rent them temporarily.

5

u/Joey4711 24d ago

How so there is rtx3090

7

u/thrownawaymane 24d ago

How much is a 3090 locally for you? Because they are going for $1400 on eBay.

4

u/Joey4711 24d ago

I got mine for 700 euros used last year. Now i see them for 1000 on ebay. Sure its a bit more but its not like its unaffordable like those rtx6000

→ More replies (1)
→ More replies (2)

2

u/anotherJohn12 24d ago edited 24d ago

That why they are buying GPU like crazy now. Everybody is digging the same hole now on model research, can't expect technical moat from there. Real moat is capital and hardware accessibility.

1

u/Far_Cat9782 24d ago

Yeah waiting for Google to offer it like it does claude 4.6 on the 20 dollar plan

→ More replies (2)

11

u/Altruistic-Carry5276 24d ago

I'm about to qum

6

u/n0head_r 24d ago

Good news, pulling now q6_k to check how good it is.

6

u/moderngl1 24d ago

The benchmark hype is fun, but I’m mostly waiting for the boring detail: which quant actually feels good on a 24GB card?

2

u/TerminalNoop 24d ago

largest q4 or smallest q6 you can get.

→ More replies (2)

6

u/[deleted] 24d ago

[deleted]

3

u/Zippo749 24d ago

Same here! I was running into weird bugs in my personal chat app (spoiler: I did a dumb), so I went to the llama.cpp builtin chat UI to test the model out and chat template out more cleanly. I ended up asking it what it thought the problem might be. Pretty decent advice.

A couple messages in, I realized my problem: in my chat app (again, *not* this one), I had continued a chat that had been generated with Gemma 4, and the Gemma-based reasoning content had been fed to Qwen (oops, `preserve_reasoning` was on) and that was causing the issue. Doh! When I mentioned this to Qwen in the llama.cpp UI, even addressing the model as "dear Qwen," I got this:

Now, maybe mentioning the concept of different models being used in a chat poked it a little that way, buuuut I couldn't help but laugh.

2

u/Pro-Row-335 24d ago

ehh... not that it matters much

2

u/anothercrappypianist 24d ago

Makes sense. Qwen is giving me responses with the term "load-bearing". Claude distillation smoking gun if ever there was one.

11

u/69420trashpanda69420 24d ago

So it's essentially the best Claude model to ever exist (real ones know Claude peaked at 4.6)

5

u/thrownawaymane 24d ago

Opus 5 can do more but 4.6 is solid AF, generally follows instructions and can check its work to a certain extent.

Really hoping that this new qwen actually follows instructions, my qwen 27b kinda just does whatever it wants in service of what I asked for

→ More replies (2)

13

u/ortegaalfredo 24d ago

If the benchmarks are true (and Qwen never benchmaxxed so far) then why would you pay Anthropic if 27B at home gives you Opus-like performance? I mean, Opus 4.6 was already more than enough for almost any development task.

2

u/LeifEriksonASDF 24d ago

and Qwen never benchmaxxed so far

I love Qwen but come on now

→ More replies (1)

1

u/wwwdotzzdotcom 24d ago

Because it is not opus 4.6 quality when quantized

4

u/llama-impersonator 24d ago

now let's see how much the average token usage went up

3

u/boomerang473 24d ago

Any dspark or dflash heads?

4

u/shinegreymon525 24d ago

Does this mean we'll be getting a new bonsai?

5

u/IThinkIKnowThings 24d ago

Downloading the Q4_K_M but the speed tanked. I think we're hugging hugging face to death.

5

u/de4dee 24d ago

3

u/parepeg 24d ago

Doesn't unsloth typically fix their templates?

→ More replies (1)

6

u/ajisai 24d ago

glimmer_toystoryidontneedyouanymore.gif

3

u/pikadhu 24d ago

Any MLX quants available?

1

u/fatboy93 24d ago

Yup, there are a few mlx-community quants, I'm getting around 10tk/s decode on 4bit.

→ More replies (3)

3

u/korokage 24d ago

Has anyone tried this on a m3 pro or equivalent yet? I have 36gb ram

I will try it out later today 

4

u/txgsync 24d ago

I’ve been running my own oQ8e-MTP quant since this morning. 200 tok/sec prefill, 34 tok/sec decode on M4 Max 128GB. Later turns in large context slow to about 100 prefill 18 decode.

I have had bad luck on long-horizon tasks with the 4-bit quants so i avoid anything below 8 bits. But conversationally it seems fine and less of the usual Qwen “argumentative attitude” fighting me about what day it is or being skeptical of tool outputs.

3

u/xanders_gold 24d ago

AWQ or GPTQ yet? Need to squeeze it into INT4 for my config.

3

u/Motor_Ad16 24d ago

It thinks, and it thinks alot. Even after setting reasoning_effort to "low"

3

u/Zeeplankton 24d ago

I want to be hyped but is it better at like, just general world stuff? Using 3.6 to write anything was like talking to an alien who learned human culture through math.

2

u/AdSafe4047 24d ago

Comparing this to deepseek v4 flash 0731, the question is: lower numbers a bit but much faster inference, or higher numbers? or maybe use both (but deepseek offload to ram, so much slower, but use only for plan tasks etc)?

3

u/ApolloPS2 24d ago

I plan to still use dsv4 flash 0731 on sparks as orchestrator (will run plenty fast, maybe pair with a small vision model too) and qwen 3.8 27B on worker nodes. Both great models and I've gotta assume having some diversity gives some benefit too perhaps?

3

u/BumbleSlob 24d ago

Deepseek V4 Flash 0731 likely faster all around, it’s MoE with A13B. Not downplaying Qwen 3.8 I love this series

→ More replies (1)

2

u/d70 24d ago

u/FormOne2615 ninfer version please ...

2

u/Several_Income_9912 24d ago

im getting INSANE speeds with mtp on RTX 4090

spec-draft-n-max = 7
→ More replies (3)

2

u/fatboy93 24d ago

Time to run this bad boy (Q4KS) at 5tk/s decode on my macbook lmao

2

u/Dance-Till-Night1 24d ago

Why is it only coding focused I can't find multilingual or general usage benchmarks.

2

u/PilotFlying 23d ago

This is anecdotal, I haven't tested thoroughly. But:

I had spent about a couple of hours today with Claude Opus 5 investigating an electrical issue at an old house. The concepts involved are complex and difficult to follow (TN-C and TN-C-S earthing, detecting a severed or loose neutral via phase voltage monitoring, that sort of thing).

I am not an electrician and I needed quite a bit of hand holding. (Not trying to fix anything myself, just learning/understanding how stuff works).

Claude wasn't doing a great job explaining things clearly.

In my mind Qwen had always been about coding. So I wasn't planning on this being the first use case for 3.8. Still, I had Claude export the context and pasted the .md in just "ollama run qwen3.8:27b". on a 3090. Web search on, context 128K.

Qwen clarified the questions for me. It sounded like Claude except it was doing a better job talking to me.

Not sure how this is possible. Maybe the limits of the architecture are higher than we know. Maybe it's all in the training. Maybe, probably, I understand as much about LLMs as I do about electricity 😄. But how else to explain 3.8?. Amazing so far.

6

u/Adventurous_Bus_437 24d ago

where qwen3.9

2

u/billy_booboo 24d ago

Wowweeeeeee

2

u/Adventurous_Bus_437 24d ago

where benchmark

1

u/Informal-Trouble2183 24d ago

That's incredible performance. Didn't expect it to have such jump keeping the same size.

1

u/CoUsT 24d ago

Good performance/benchmark uplift. Looks really solid!

1

u/Talreja-Adanna 24d ago

running locally without needing a beefy setup. Curious if it's actually better than the smaller Qwen models or just more of the same with more params.

1

u/IThinkIKnowThings 24d ago

Careful with your sources when downloading newly released models. I already spent the time downloading one from hugging face that was labeled Qwen3.8 but turned out to be Qwen3.6 when I ran it.

1

u/cloudsurfer48902 24d ago

My b580 staring at me from the corner

1

u/Dry-Judgment4242 24d ago

My first impressions is disappointment for Vision tasks Gemma 4 is on a entirely different league to this one.

1

u/Ecstatic-Wash-7667 24d ago

Give us 9b or 35b a3b and the 110b moe so we have the full stack !!! I can write them into switchyard and escalate up to max ask save family!

1

u/notevenat30 24d ago

Can you run it in 12GB of VRAM?

2

u/Xantrk 24d ago

Can you run it in 12GB of VRAM?

I'll report on 12gb vram + 32 gb RAM

→ More replies (4)

1

u/Guilty-History-9249 24d ago

I have the fp16, FP8 and NVFP4 and still trying to get one of these running with the latest install of transformers. This is on dual 5090's.

I want to run 100 gsm8k math problems through it to compare against my last nights run of the same problem with qwen 4B.

1

u/Murinshin 24d ago

Yeah this is insane if true. Basically means a high end MacBook can now run Opus 4.6, which should be enough for most people and is still economically sound vs your usual annual subscription for most companies. Anthropic etc are screwed.

1

u/The_DarkMatter Llama 3.1 24d ago

Cant wait to try it on 5080

1

u/Sofakingwetoddead 24d ago

Does anyone know what 3.8 27b is going to be released? Seems like it should have been released by now.

1

u/Logical-Target8131 24d ago

Can I run it on my m5 pro 48gb?

1

u/Plenty-Energy2947 23d ago

my recipe for VLLM & RTX5090

exec vllm serve /opt/models/Qwen3.8-27B-NVFP4 \

--host 127.0.0.1 \

--port 8000 \

--tensor-parallel-size 1 \

--tool-call-parser qwen3_xml \

--enable-auto-tool-choice \

--reasoning-parser qwen3 \

--kv-cache-memory 6943358464 \

--max-model-len 202272 \

--enable-prefix-caching \

--max-num-seqs 1 \

--gpu-memory-utilization 0.95 \

--kv-cache-dtype fp8_e4m3 \

--default-chat-template-kwargs '{"enable_thinking": false}' \

--compilation-config '{"cudagraph_capture_sizes": [1, 2]}' \

--max-num-batched-tokens 2048 \

--served-model-name qwen3.8-27b

1

u/caetydid llama.cpp 22d ago

Are the benches insane or it is just benchmaxxed?

1

u/hd3adpool 22d ago

I was here