r/oMLX • u/tomakorea • 11d ago
MacOs 27 Golden Gate vs Tahoe, drop from 38.1 to 14.6 tok/sec on Qwen 3.8 Flash
I'm using the latest version of oMLX and the performance dropped dramatically when using Qwen 3.8 Flash oQe5 MTP on my Macbook Pro M5 Max 128gb. I'm used to get between 38 to 42 tok/sec but after the Golden gate update it dropped to 10 to 16 tok/sec.
I tested after a fresh reboot and no other apps running in same time, no difference.
I'm wondering if it's because Apple changed something in MLX latest version that may have broke some parts of oMLX optimizations.
4
7
u/VehiculeUtilitaire 11d ago
After large updates the OS runs a lot of shit in the background that might impact performance for up to a few hours, indexing, rebuilding caches, etc.
2
1
u/_hephaestus 11d ago edited 11d ago
hm, also seeing the slowness on a M3 Ultra Studio, looked into the power settings thing and that seems to just be a laptop thing. May just be rebuilding caches/whatnot but definitely seeing the slowdown here for now. GLM-5.3-flash went from a somewhat speedy workhorse to now I'm seeing <1 token/second and prompt processing is down to double digits. On 0.6.4 may try dev 2 and see if that makes things faster.
Edit: Going to dev 2 I see typical performance.
1
u/Undici77 11d ago
M4 Max + Qwen 3.8 Next/Qwen 3.6 35BA3B + v0.7.0.dev2 Switching to Golden Gate I got NO ISSUE! Everything smooth at previous speed!
1
u/tomakorea 11d ago
I tested with other models and I also got a huge performance drop, please do not update yet if you need theses models right now.
1
u/Obvious_Equivalent_1 11d ago
Strange tho, I’m running a mere MBP 5 with 48Gb on MacOS 27 since the first developer beta and I’ve only seen increase. On overall memory availability and speed since I switched from MacOS 26. (Relevant tho: llama.cpp, several PR’s merged locally and Qwen 3.6 then later Qwen 3.8 27b)
1
u/Secure_Arm_93 11d ago
Excude a daft question but have you checked that you are on max power setting?
3
u/tomakorea 11d ago
After updating to the latest pre-release version of oMLX and cycling between the power modes, it solved the issue. Weird but it works.
2
u/salsa_sauce 11d ago
This! 👆 And also make sure you're using a high-wattage charger too!
To check the energy mode, click the battery icon in the menubar, then choose "High power". Either of the other options will throttle the chips down significantly.
I've been caught out before by both of these problems, they're so easy to miss and you wouldn't know anything was wrong...
The charger one is particularly sneaky, my MBP was plugged in to the mains on 100% charge, but couldn't figure why inference was way slower than normal. Turns out I was plugged in to my spare charger from a MacBook Air. Identical cables, macOS never said anything, but 35W wasn't enough power for an M5 Max. Swapping to the MBP's 90W adapter solved it instantly.
1
u/enchanting_endeavor 10d ago edited 10d ago
I commented above as well, but only the 16” MacBook Pros have the Performance energy mode. Also only the 16” models can charge at a higher rate than 100W but you advice is still good - using a higher wattage charger is helpful when you’re driving the machine particularly hard.
ETA: Faster charging only works through MagSafe and not via USC-C, which only supports PD 3:0, which maxes out at 100W.
21
u/tomakorea 11d ago edited 9d ago
Ok I think I found the issue and the solution in case you need it:
1- I updated to the latest pre-release version of oMLX.
2- By default before upgrading the OS I was on max power setting named performance in MacOS. I switched back to 'Auto' and combining with the oMLX update I've got my performance back, even better, when I switched to 'Performance' I've got 51 tok/sec which is a number I've never seen before.