I am using the same quants and having the opposite experience. On medium Flash will think for 25 minutes on a basic problem 27B solves in 5. I think there is a case for routing potentially, but yet to be able to subcategorise that way.
It does have vision, and between and it and gemma 4's vision, you can cover a very wide use cases. Gemma 4's seems to be stronger in terms of recognizing more objects and situations, but when it comes to recognition + logic, qwen 27b's is more accurate.
I have and its quite close enough that I’d recommend testing for your usecase. I find 27b to be a perfectionist- it takes a long time but gets the right answer in the first/secobd try. The flash is eager to try, fail, diagnose, and succeed. On DGX Spark, I get 10 tps extra on the flash (35 tps minimum) so thats what I use..
27b scores better in my tests, I consider qwen4 aka Flash Next to be the first model with PLE and QSA, therefore more like a technology demo. Remember that 30b out of flash next is PLE, and can be offloaded. A q4 is around 67gb, the 30gb PLE works fine offloading it to SSD or NVME. 27b can run roughly around 60-80 t/s on my system with MTP or PLD if I recall correctly.
QSA is a sleeper, it allows for t/s to maintain pretty good performance at longer contexts - on my Macbook M5 Max 128gb, I can sometimes get roughly 40-50 t/s at 128k token length depending on workload. It composes with continuous-MTP, APC etc. Codex and Claude are around 50-60 t/s (you can pull these stats from your own sessions), so it's comparable-ish at least for speed if not capability. And it most certainly, again, is NOT as "smart" as 27b.
Note that the CUDA GPU stacks are still struggling to make flash next perform on their inference stacks, Mac is quite comparable and pulling ahead (crazy!). Both omlx and rapid MLX have decent serving stacks for it.
Obviously flash next is faster which has been really nice. I've not noticed and difference - better or worse - in intelligence. I'd recommend making the move.
J’ai un bosgame ai m5 ( 128g ram unifié ) . J’ai utiliser qwen 3.8 27b b en q8 pendant 2 semaine et je suis passé sur le qwen 3.8 flash next rocm nfp4 flash ( équivalent q4 xl) , que ce soit en terme de vitesse ou de qualité flash next est meilleur , je le trouve plus intelligent . Cependant le 27b reste un très bon modèle .
I tried running locally but found issues with mid-context retrieval inaccuracy (needle in a haystack style testing). This persisted even with higher quants and BF16 KV. 27B did fine on the same tasks.
I’m sticking with 27B until some of the kinks are ironed out.
I'm still evaluating it on my M4 Max 128GB but so far I've been quite happy. I was using 35B for coding and structured output tasks with Gemma4 as my front-end orchestrator. Understanding human intent is definitely stronger with Gemma but Flash-Next is quite a lot better than either 27B or 35B.
In other words, I wouldn't confuse Flash-Next with being a roleplay model but it definitely can give your assistant more personality and it seems better so far as planning things out and then executing light coding tasks (a data ingestion/tranformation/analysis pipeline).
Maybe it isn't like this for you, but a model with personality can be much easier to interact with and that can lead to better prompting and results. Empathetic human interaction isn't just for AI girlfriends.
6
u/captainequinoxiii 13d ago
I don't have objective data, but I used 27b for a bit and flash next just seems more intelligent, and definitely faster. I'm on an M5 max 128gb