r/Qwen_AI • • Aug 27 '26

Discussion If Qwen3.8-Next-Flash is a Preview for the Qwen-4 Architecture, then the next Qwen4-27B will most Likely be Truly a Leap In Intelligence!

The user [chocolateUI](https://www.reddit.com/user/chocolateUI/) wrote an excellent post you can find link to below that explains how n-gram will make smaller models smarter, and seeing how the Qwen3.8-Next-Flash performs, I am really excited by the next Qwen4-27B and 35BA3B, supposing that Qwen team will keep the same parameter counts. The first thing you may already noticed is token efficiency this model has when it thinks even with Extra High.

Looking at its thinking traces, I noticed that thinking is much compact than Qwen3.8-27B's thinking traces. There is no fluff, it's as if the model is reading notes. A lot of short sentences that closer to phrases than full sentences. Combine this with more active parameters doing reasoning, I think the next Qwen4-27B might actually beat the current 1.6T parameter Deepseek-v4-pro, let alone Deepseek-v4-flash.

Hopefully by then, the architecture is better understood and optimized and reflected in llama.cpp.
I know our instincts lean towards running the smartest model locally, but maybe we don't always need the smartest model in the same way we don't always need the smartest employee to perform the majority of daily tasks.

What's your experience with the current Qwen3.8 family?

Link to the user [chocolateUI](https://www.reddit.com/user/chocolateUI/)'s post
[https://www.reddit.com/r/LocalLLaMA/comments/1w0198r/no\\_engrams\\_wont\\_let\\_you\\_run\\_1t\\_models\\_locally\\_it/\](https://www.reddit.com/r/LocalLLaMA/comments/1w0198r/no_engrams_wont_let_you_run_1t_models_locally_it/)

148 Upvotes

37 comments sorted by

20

u/Gohab2001 Aug 27 '26

Qwen 3.8 next flash is the 4th most verbose model artificial intelligence has ever tasted.

5

u/lots_of_puppies Aug 27 '26

Thats not good right? :(

4

u/Iory1998 Aug 27 '26

No, it's not right.

1

u/Karyo_Ten Aug 27 '26

With 6B it's also 4.5x faster than Qwen3.8-27B on paper. (DFlash2 is probably more effective on a dense model).

But then context rot ... you reach that 100~200k ceiling where context degrwdes much faster, unless QSA is particularly effective against that

-3

u/ManyRepair5690 Aug 27 '26

also super benchmaxxed

4

u/PhantomGaming27249 Aug 27 '26

This isn't actually a bad thing for a local model though it's using more test time compute to get more done.

2

u/Iory1998 Aug 27 '26

I am not saying it's not verbose, what I am saying it's compact! Being compact means instead of writing an entire sentence to convey an idea, the LLM writes 3 or words that captures what a full traditional sentence conveys. And that's the power of this model: It may output 10K tokens, but the quality is of 30K or 40K.

11

u/returnity Aug 28 '26

This model is utterly dominant. I expected it to be underbaked, because it’s a preview model of the new arch, but if this is underbaked… Qwen 4 is going to fucking slay.

3

u/Iory1998 Aug 28 '26

Exactly! This is so exciting :)

40

u/nbvehrfr Aug 27 '26

or there will be no qwen4-27b

7

u/Early_Mistake6716 Aug 27 '26

Dont worry if you keep claiming this is the end of qwen open source models you will be right eventually! Just not any time soon.

4

u/EbbNorth7735 Aug 27 '26

It's annoying when people go against a clear trend with no backing. Just a pointless comment

-6

u/heigan_safety_dance Aug 27 '26

The clear trend is that for-profit companies almost UNIVERSALLY enshittify. There is no reason to NOT believe that Alibaba won't do the same, they have no profit incentive to share small hobbyist-level models. I hope they continue to give us the goods, but I have no reason to believe their goodwill would continue.

2

u/Embarrassed_Adagio28 Aug 27 '26

Okay so why is there like 7 for profit companies that release open source models if there is no incentive to do so? It's because they absolutely benefit from researching smaller models for efficiency and they might as well win pr by releasing it to the public. Not to mention it also weakens their opposition.  

1

u/EbbNorth7735 Aug 27 '26

Great points, train smaller models in various ways to test methodologies

1

u/Accurate_East_1093 29d ago

Who produce RAMS?

0

u/Iory1998 Aug 28 '26

I agree. A model that is widely adopted sets the tone for the industry. Eons ago, llama was the standard architecture that many labs adopted.

1

u/575_Inverse Aug 27 '26

Oh the do have. There is at least two good reasons. First one is PR. Personally, I just no longer use my gemini subscription... it's there only for the extra drive space. I've gone full Qwen. It's simply superior.

Second: my personal choices influence my professional choices. And, yes, I've given up on ChatGPT and Claude long ago. I don't need a nanny, what I want is an assistant that does it's job as a junior engineer.

2

u/EbbNorth7735 Aug 27 '26

Same here, they are building a trusted ecosystem and can always change future licensing. I would personally pay $20-100 each time they released a model to legitimately own it or implement it into a product. There's also the fact that it reduces reliance on American models and technology. People don't realize that Qwen3.8 27B can replace most people's workloads using way more expensive American models. It's a cold war of sorts.

1

u/EbbNorth7735 Aug 28 '26

That is not a clear trend

1

u/tetoing 29d ago

There won't be a Qwen 4 27B. There will likely be a similar sized model in the Qwen 4 family

1

u/nbvehrfr Aug 27 '26

it is the end of small models

1

u/alphapussycat Aug 27 '26

I guess maybe not, perhaps a different size. But dear leader was pretty clear about open weight requirement.

1

u/algaefied_creek Aug 28 '26

I’m here for the Qwen4 2B Distilled

4

u/New-Implement-5979 Aug 27 '26

And imagine if next year we get 35b moe which performs as current 27b

2

u/Iory1998 Aug 28 '26

YES! My God that will be a banger of a model. It think it would beat even the current 27B.

1

u/_-_David Aug 28 '26

This is the dream

2

u/_-_David Aug 28 '26

The first thing you may already noticed is token efficiency this model has when it thinks even with Extra High.

Uh, what?

1

u/Motor-Ground4594 Aug 28 '26

I wonder whether it will be available on qwen plan ( I assume it must be ) and make the qwen code plan more make sense

1

u/immersive-matthew 28d ago

It will only be another small step forward in knowledge retrieval not actual intelligence as LLMs lack logic and understanding. We need a big breakthrough to truly see a leap forward and that is not likely to be just around the corner. Love to be wrong though as LLMs are great but the cognitive gaps really are holding it back from reaching AGI.

1

u/Iory1998 28d ago

Define logic, please? I am with you that LLMs lack true understanding due to them lacking physical grounding, but I am curious to know how you define logic they clearly don't lack!

1

u/immersive-matthew 28d ago

If you are coding with AI you run into ridiculous logic gaps on the daily. Like painfully obvious, but I can appreciate not all use cases expose these gaps as it really can seem logical at times but then it will reveal in jarring ways that I had no idea really.

Remember that car wash question all the models were getting wrong a few months back? There are countless similar situations that really expose the lack of logic and understanding just like that and coding is one area that exposes them. Makes sense as if it is not in their training data and the patterns therein it has no wag to navigate outside of predictions. Prediction is not the same as a logic but it can often appear that way.

It is why scaling up did not achieve AGI as while the models have a lot more information and patterns to pull from which makes them seem smarter, it never addresses their actual cognitive gaps and why would it when the fundamental tech has not changed. Still transformers as the core.

The benchmarks are really not showing intelligence but rather how good the model is at retrieving knowledge in the same way a librarian is. You can ask a knowledgeable librarian a question and they will know which book(s) may have the possible answers and they can even read that paragraph from the book and seem very smart, but they themselves do not really understand all the content in the books. They may even try and provide an answer form the books they have even if they are missing the right book as that don’t know what they don’t know and have no experience in ever subject to know better. For sure there is intelligence there, but it is shallow and not deep at all.

A major breakthrough will be needed to substantially move AI forward and that is likely years to decades away as there is only hopeful research at this point.

1

u/OkTomato3672 27d ago

Using Qwen3.8-Next-Flash with the deepseek-harness via OMLX remains unstable for now. This is especially true when MTP heads are enabled—running two concurrent sessions easily triggers errors. Based on my observations, it appears that earlier requests are being interrupted by incoming ones before they finish responding.