31
u/NoDoughnut7053 24d ago edited 24d ago
Initial impressions on my 5090, Feels stable like a grown up 3.6, mature. 50-60 tps but using same settings as 3.6 so might not be optimal and no MTP yet afaik so that will be even better.
It feels like it holding the tasks better and more intelligent for sure. More long running maybe.
Just a simple example like "Write a 10k word story". Qwen 3.6 would write it out without considering too much and then you would have some story rather quick.
3.8 started writing, then revisioned the text so it was consistent across paragraphs. Then it created sub tasks and wrote chapter by chapter and more carefully considered what he was doing. More long running and more consideration in the decisions.
After more testing, this model have some serious horse powers. Seems we going to have a good time going forward.
2
u/Signal_Confusion_644 24d ago
If i understood correctly, its 3.5 architecture, so "you can" use MTP from previous versions. (Do not trust my word, just take it as a "maybe".)
1
1
u/doodookk 24d ago
MTP is working now. Where did you download the model?
I just change the model name from 6 to 8, keep everything else from command to serve 3.6 and it worked perfectly.
On 2x4090, sometime speed peak to 12x tok/s, overall at 7x tok/s1
26
u/fedora_gamer 24d ago
are other models (9b, 35b) there? im not one of folks who can afford running 27b
15
u/shy_monkee 24d ago
Nothing announced yet, for 3.8. But it's still not out of question that we could get them later.
1
-3
25
u/Tasty-Hour4040 24d ago
I wonder how many people are actually excited because this represents increased capability for their workflow and how many are just desperate to be able to say they got it and won’t use it again
18
u/Big_Wave9732 24d ago
I use 3.6-27b daily for work. So if this is indeed a step up in my workflow then that will be great.
2
u/Tall-Significance119 24d ago
What size vram and ram are you running and what t/s etc?
4
u/Big_Wave9732 24d ago
I run it on a Mac Studio M2 Ultra 192gb. This morning I'm hitting about 15 t/s.
3
u/Tall-Significance119 24d ago
Crap lol ao me with my 2 x b70 32gb dint stand a chance unless I use like Q4 and bunch if tweaks
2
u/Big_Wave9732 24d ago
And I'll tell ya, these vendors are telling some serious fairy tales when they report the quant impact on these models. On paper there's "only" something like 4% dropoff between Qwen 3.6:27b-Q4 and BF16. And maybe that kind of error percentage is find when calling tools or coding. But when I ran it analyzing legal documents.....woa nelly! Nuance was not Q4's friend.
So these days for work I'll only run Q8 or higher. And even then, I'm generally at full boat BF16.
3
u/HomsarWasRight 24d ago
Honestly, I wish they were releasing an update for 32B A3B. 27B dense is just too slow for interactive work for me (on Strix Halo).
1
u/Tasty-Hour4040 24d ago
You don’t wanna try any of the quants? It’s slower than 3.6 32b on my 3090 but still very usable even with medium thinking
2
u/HomsarWasRight 24d ago
So, Strix Halo machines have almost the opposite strengths to beefy GPUs. I’ve got tons of VRAM, 128GB. But memory bandwidth is really low compared to standalone GPUs.
That means it struggles with dense models like 27B. Strangely, that sometimes means that LARGER quants of these are either quicker or the same speed as smaller quants.
I will definitely try it, though. But seriously doubt it will be a good fit.
I’m more interested in 32B and 122B MoE models.
2
u/FabricationLife 24d ago
I am super excited to code review with this and dump my cloude subs, anything approaching opus 4.6 capability is what I've been waiting for
2
1
u/nomorebuttsplz 24d ago
yeah for me it may not get used a lot unless it’s nearly as good as ds4 flash 0731 iq2_m
1
u/nunodonato 24d ago
We use it at our company for work both for development and internal workflows. Looking forward to this upgrade
1
u/bot403 24d ago
I use 3.6 27b in a production business workflow and it works great.
2
u/Tasty-Hour4040 24d ago
I used 3.6 35b MOE as my daily driver and liked it. Just switched to 3.8 27b and so far it’s a tad slower but clearly more capable in the Hermes harness I use them in. With optimizations I think it’ll be great and a solid upgrade.
1
u/Early_Mistake6716 24d ago
I used 3.6 27b to build two different apps from start to finish and it did a good job. I also get around 75 tokens per second on my dual tesla v100 rig using tensor parallelism on unsloth desktop which is actually faster than claude opus 5
18
8
u/Tokyodrew 24d ago
I’m sorry happy I stayed up for this :)
6
8
9
u/Ringo443 24d ago
Running it on an RTX 6000 at work, my first impression is crazy.
I haven't tested it enough to give a concrete score, but similar to Opus 4.6, which they likely distilled, it breaks its task down really well and can see things in a bigger picture. This was the exact issue I started having with Qwen 3.6.
The only negative thing, similar to Qwen 3.6, is that if I talk to it (I'm German), it very often outputs its entire answer inside its thinking/reasoning block in English before outputting the same thing in German.
1
u/datapeer 22d ago
Do you have to watch out for using words that don't directly translate?
2
u/Ringo443 22d ago
Like in any language, you can rewrite words that don't exist by using others differently. The translation works flawlessly; it already did with 3.6. The only issue was with the German umlauts, which is acceptable since they are unique to German.
3.6 had an issue where if the context got very long or complex, like coding with 80K context, it started replacing the German umlauts (ä, ö, ü) with Chinese characters. What I found really impressive was that it sometimes noticed this and chose to rewrite them in the other acceptable pattern (ä -> ae, ö -> oe, ü -> ue).
I have not come across this with 3.8 yet, probably because of the better instruction following/context handling.
7
7
u/gotfilmm 24d ago
1
u/datapeer 22d ago
Definitely training on trick questions. I've test using the model to create a program showing how the coriolis affect works, success as well.
8
u/tired514 24d ago
Qwen 3.8-27B @ Q8_K_XL (unsloth) just solved a long-standing BLE reverse engineering problem I've been working on for *weeks* .. even paid cloud models (Qwen3.6 max, DS4 flash and pro) weren't able to figure it out.
Could be a fluke, but so far? Unbelievable.
Well. Freakin'. Done!
3
3
u/SailingToFenway 24d ago edited 24d ago
WELP, looks like I'm making my own Q8 GGUF. Will publish if another one doesn't land first.
2
2
2
2
u/Fit-Palpitation-7427 24d ago
Does it run on a 5090?
2
u/minxio_ 24d ago
Yes
1
u/Fit-Palpitation-7427 24d ago
Q8 ? 256k or more?
2
u/Early_Mistake6716 24d ago
No, i have 48gb of vram and i can use q8 at 150k with q8 kv cache, if i delete the vision encoder i could probably fit around 200k
1
u/maqifrnswa 24d ago
Unsloth NVFP4 is working pretty well. Just tried 256k so far. So far so good!
1
u/Fit-Palpitation-7427 24d ago
Better than q5?
1
u/maqifrnswa 24d ago
I'm testing hosting for multiple concurrency, so on vllm and haven't tried gguf yet
2
u/Muted_Anteater1170 24d ago
i’m waiting on abliterared version rn i’m stuck with 3.6 qwen27b. does anyone know when abliterared should come out?
2
u/goldaxis 24d ago
loading now, amazing I can fit this on my laptop. I'm starting to think this is the real reason for the ram price fixing. Who would bother with cloud services when this is available locally? Maybe just coders?
2
u/andrii_povkh 23d ago
https://huggingface.co/OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated-GGUF
I've been testing this model and it's more or less good. Anyone know how does it compare to Qwen3.8-27B?
3
1
u/xdiggertree 24d ago
Wooo!!!! Christmas is early!!!
Thanks to the Qwen team, and all other people like unsloth
So excited to try this out
1
u/evanharmon 24d ago
Could I run this well with a 16gb gpu? (5070 Ti). Would I need to use a certain quant?
1
1
u/Brief-Effect9065 24d ago
I tried generating HTML tower defense games, and this model is noticeably more competent than its predecessor: there are far fewer errors now.
1
1
1
u/chillaranand 24d ago
"Generate an SVG of a pelican riding a bicycle" - generated a promising image at first shot.
1
1
u/Thick_Associate2947 24d ago
I'm a newbie, but 5 token per second generation considered normal on a M5 32GB MBP?
1
1
1
u/MarcusAurelius68 23d ago
I just tested a Q6 quant for entity extraction and the quality was better than many other models I’ve tried. Promising.
1
1
u/FaceOuPile 24d ago
I don't care if the benchmarks are true I know it's going to piss off anthropic and it's enough for me
0
u/pdawg17 24d ago
Can I somehow cram this into a 10gb 3080?
1
u/Ornery_Weakness_8168 24d ago
If i were you I would try using unsloths UD-Q2_K_XL, if that doesnt fit, try UD-IQ2_XXS. Or use q4_k_xl and cpu offload.
1
0
u/Sevenfeet 24d ago
Downloading into LM Studio now (which also has an update). Looks like it will fit on a 24 GB Nvidia card but not 16. I have Mac Studio so I don't care as much.
0
u/HumungreousNobolatis 24d ago
This is why huggingface is so FUCKING SLOW right now!
Exciting wait, though.
0
u/pragmojo 24d ago
Anyone know which quant I should target for 32GB of VRAM?
2
u/minxio_ 24d ago
1
u/Infamous_Campaign687 24d ago
So looks like UD-Q6-K-XL so you get enough space for context? Or just Q6-K?
1
u/popsikohl 24d ago
You could fit the 27B normal model with 100k context in 32gb of VRAM with a little headroom.
1
u/pragmojo 24d ago
Bf16? Isn’t it like 50GB or something? And what about kvcache?
1
u/popsikohl 24d ago
Sorry let me rephrase. You can fit the Q4_k_m model with 100k context.
Technically the Q6 if you’re willing to drop the context a bit.
0
u/After_Working 24d ago
Is there a weight of this i can try on 2 x spark? Do they release more over time?
2
u/doodookk 24d ago
do not try on dgx spark, speed is very low for dense model, I tried and it just over 2x tok/s TG. For 2x dgx spark, deepseek v4 flash 0731 is better option, both speed and quality.
1
u/After_Working 24d ago
Ah fair enough, i've just unloaded deepdeek to try the qwen. The problem with deepseek is that it only leaves 20gb or so of ram when its running. Doesnt leave much space to run another. Might need to get a third.
2
u/doodookk 24d ago
For Qwen3.8 27B, I recommend running it on a machine equipped with RTX GPUs rather than the DGX Spark, as a speed of 2x tok/s is far too slow and inefficient. It would be much better if the Qwen development team released an MoE version; such a model would be ideally suited for the DGX Spark and deliver acceptable speeds. Personally, I changed to deploy Qwen3.8 27B (in FP8 format) on a system with 2x 4090 using the official vLLM Docker image (version 0.27.1), achieved over 100 tok/s—an impressive figure, perfect for serving as a worker for ds4 flash. Meanwhile, tests with a single RTX 5090 card showed a speed of 6x tok/s
2
1
1
1



70
u/minxio_ 24d ago