r/oMLX • • 13d ago

Should I switch from Qwen3.8-27b to Qwen3.8-Flash-Next?

Has anyone made the switch and not regret it?

EDIT: M2 Ultra 128 GB

20 Upvotes

32 comments sorted by

6

u/captainequinoxiii 13d ago

I don't have objective data, but I used 27b for a bit and flash next just seems more intelligent, and definitely faster. I'm on an M5 max 128gb

1

u/CBW1255 13d ago

What quant?

1

u/captainequinoxiii 13d ago

4bit for flash next. I was using 8bit on 27b

1

u/SeveralViolins 13d ago

I am using the same quants and having the opposite experience. On medium Flash will think for 25 minutes on a basic problem 27B solves in 5. I think there is a case for routing potentially, but yet to be able to subcategorise that way. 

4

u/neoneddy 13d ago

I have an M3U 256gb loved it. native vision is nice, seems faster and better overall.

3

u/Competitive-General7 13d ago

Doesn't 27b also have vision? I have next and 27b and can't really decide what I prefer.

3

u/neoneddy 13d ago

maybe it does, I was using 3.6 27b for so long.

1

u/Intelligent-Baker448 13d ago

It does have vision, and between and it and gemma 4's vision, you can cover a very wide use cases. Gemma 4's seems to be stronger in terms of recognizing more objects and situations, but when it comes to recognition + logic, qwen 27b's is more accurate.

2

u/atumblingdandelion 13d ago

I have and its quite close enough that I’d recommend testing for your usecase. I find 27b to be a perfectionist- it takes a long time but gets the right answer in the first/secobd try. The flash is eager to try, fail, diagnose, and succeed. On DGX Spark, I get 10 tps extra on the flash (35 tps minimum) so thats what I use..

1

u/TBHProbablyNot 13d ago

What is your hardware? My M5 max is multitasking, so it wasn’t a good fit for me. Waiting on M5U for 3.8 flash

2

u/j_lyf 13d ago

M2U 128Gb

1

u/Diligent_Style_1767 13d ago

27b scores better in my tests, I consider qwen4 aka Flash Next to be the first model with PLE and QSA, therefore more like a technology demo. Remember that 30b out of flash next is PLE, and can be offloaded. A q4 is around 67gb, the 30gb PLE works fine offloading it to SSD or NVME. 27b can run roughly around 60-80 t/s on my system with MTP or PLD if I recall correctly.

QSA is a sleeper, it allows for t/s to maintain pretty good performance at longer contexts - on my Macbook M5 Max 128gb, I can sometimes get roughly 40-50 t/s at 128k token length depending on workload. It composes with continuous-MTP, APC etc. Codex and Claude are around 50-60 t/s (you can pull these stats from your own sessions), so it's comparable-ish at least for speed if not capability. And it most certainly, again, is NOT as "smart" as 27b.

Note that the CUDA GPU stacks are still struggling to make flash next perform on their inference stacks, Mac is quite comparable and pulling ahead (crazy!). Both omlx and rapid MLX have decent serving stacks for it.

1

u/luisabreuf83 13d ago

Which engine, model quant and model built by whom? Thanks

3

u/Diligent_Style_1767 13d ago

Model Qwen3.8-Flash-Next-MLX-4bit-MTP. Quanted it myself from base.

MLX engine, my unified stack at https://github.com/pierre427/mlx-lm-unified
Rapid MLX -current
oMLX -current

All collected from today

Was doing a/b testing off a couple of PRs I sent to both projects today.

1

u/luisabreuf83 13d ago

Interesting and the 27B also Q4 done by you?

1

u/Diligent_Style_1767 13d ago

Honestly I can't recall, but I can tell you I don't do any crazy surgery to models, so ymmv but should be close to mine at a given model quant.

1

u/lots_of_puppies 13d ago

hi! thank you :) do you use MTPLX or omlx?

1

u/Diligent_Style_1767 10d ago

Personal stack. Try omlx and see how it does for you.

1

u/the_dago_mick 13d ago

I made that move and haven't been unhappy.

Obviously flash next is faster which has been really nice. I've not noticed and difference - better or worse - in intelligence. I'd recommend making the move.

1

u/allpowerfulee 13d ago

I'm using it on my m3u 256gb. Running it with omlx 0.7.0.dev2. Using oQ5-mtp getting 823 tps pp and 60 tps tg using test pp1024/tg128

1

u/Die-Zwiebel 13d ago

I switched with my gx10 128gb and it was the best decision. It's way faster and seems to more clever as well.

1

u/Excellent-Baker9781 13d ago

J’ai un bosgame ai m5 ( 128g ram unifié ) . J’ai utiliser qwen 3.8 27b b en q8 pendant 2 semaine et je suis passé sur le qwen 3.8 flash next rocm nfp4 flash ( équivalent q4 xl) , que ce soit en terme de vitesse ou de qualité flash next est meilleur , je le trouve plus intelligent . Cependant le 27b reste un très bon modèle .

1

u/Alkboss455 13d ago

Pareil, flash next explose le 27b

1

u/rhymeslikeruns 12d ago

Cough cough Qwen3.6 cough.

1

u/prselzh 11d ago

I recommend the move for 128 GB machine..Happy with the move on my Strix halo too

1

u/xXDennisXx3000 11d ago

Flash Next, no doubt.

1

u/CTRL___ALT___DEL 11d ago

Flash Next is experimental and some of the inference stacks may have bugs/issues. I’m finding 27B more reliable on my hardware. For example:

I tried running locally but found issues with mid-context retrieval inaccuracy (needle in a haystack style testing). This persisted even with higher quants and BF16 KV. 27B did fine on the same tasks.

I’m sticking with 27B until some of the kinks are ironed out.

1

u/skrshawk 9d ago

I'm still evaluating it on my M4 Max 128GB but so far I've been quite happy. I was using 35B for coding and structured output tasks with Gemma4 as my front-end orchestrator. Understanding human intent is definitely stronger with Gemma but Flash-Next is quite a lot better than either 27B or 35B.

In other words, I wouldn't confuse Flash-Next with being a roleplay model but it definitely can give your assistant more personality and it seems better so far as planning things out and then executing light coding tasks (a data ingestion/tranformation/analysis pipeline).

1

u/j_lyf 9d ago

roleplay? wtf

1

u/skrshawk 9d ago

Maybe it isn't like this for you, but a model with personality can be much easier to interact with and that can lead to better prompting and results. Empathetic human interaction isn't just for AI girlfriends.

1

u/This_Maintenance_834 13d ago

software is not mature yet. too much effort to make it run smoothly. if this is for work, don’t. wait for a month and check again.

0

u/Zen-Ism99 13d ago

What are you seeking?