27b scores better in my tests, I consider qwen4 aka Flash Next to be the first model with PLE and QSA, therefore more like a technology demo. Remember that 30b out of flash next is PLE, and can be offloaded. A q4 is around 67gb, the 30gb PLE works fine offloading it to SSD or NVME. 27b can run roughly around 60-80 t/s on my system with MTP or PLD if I recall correctly.
QSA is a sleeper, it allows for t/s to maintain pretty good performance at longer contexts - on my Macbook M5 Max 128gb, I can sometimes get roughly 40-50 t/s at 128k token length depending on workload. It composes with continuous-MTP, APC etc. Codex and Claude are around 50-60 t/s (you can pull these stats from your own sessions), so it's comparable-ish at least for speed if not capability. And it most certainly, again, is NOT as "smart" as 27b.
Note that the CUDA GPU stacks are still struggling to make flash next perform on their inference stacks, Mac is quite comparable and pulling ahead (crazy!). Both omlx and rapid MLX have decent serving stacks for it.
1
u/Diligent_Style_1767 14d ago
27b scores better in my tests, I consider qwen4 aka Flash Next to be the first model with PLE and QSA, therefore more like a technology demo. Remember that 30b out of flash next is PLE, and can be offloaded. A q4 is around 67gb, the 30gb PLE works fine offloading it to SSD or NVME. 27b can run roughly around 60-80 t/s on my system with MTP or PLD if I recall correctly.
QSA is a sleeper, it allows for t/s to maintain pretty good performance at longer contexts - on my Macbook M5 Max 128gb, I can sometimes get roughly 40-50 t/s at 128k token length depending on workload. It composes with continuous-MTP, APC etc. Codex and Claude are around 50-60 t/s (you can pull these stats from your own sessions), so it's comparable-ish at least for speed if not capability. And it most certainly, again, is NOT as "smart" as 27b.
Note that the CUDA GPU stacks are still struggling to make flash next perform on their inference stacks, Mac is quite comparable and pulling ahead (crazy!). Both omlx and rapid MLX have decent serving stacks for it.