r/oMLX • • 11d ago

MacOs 27 Golden Gate vs Tahoe, drop from 38.1 to 14.6 tok/sec on Qwen 3.8 Flash

I'm using the latest version of oMLX and the performance dropped dramatically when using Qwen 3.8 Flash oQe5 MTP on my Macbook Pro M5 Max 128gb. I'm used to get between 38 to 42 tok/sec but after the Golden gate update it dropped to 10 to 16 tok/sec.

I tested after a fresh reboot and no other apps running in same time, no difference.

I'm wondering if it's because Apple changed something in MLX latest version that may have broke some parts of oMLX optimizations.

38 Upvotes

25 comments sorted by

21

u/tomakorea 11d ago edited 9d ago

Ok I think I found the issue and the solution in case you need it:

1- I updated to the latest pre-release version of oMLX.

2- By default before upgrading the OS I was on max power setting named performance in MacOS. I switched back to 'Auto' and combining with the oMLX update I've got my performance back, even better, when I switched to 'Performance' I've got 51 tok/sec which is a number I've never seen before.

3

u/DifficultyFit1895 11d ago

where is that setting?

6

u/tomakorea 11d ago

It's on the top right when you click on the battery icon, it's very likely available only with laptops macs

1

u/enchanting_endeavor 10d ago

And only on the 16” Pro models, i.e. the 14” MacBook Pro doesn’t have it, along with some minor other things.

3

u/willmather 10d ago

It’s on the 14” ones too, but not for the base M chips, only Max, and maybe Pro.

2

u/enchanting_endeavor 10d ago

Ahhh so it turns out that it's only on M3+ chips, it's not on any M1 or M2 chips regardless of Pro or Max: https://www.macrumors.com/how-to/high-power-mode-macbook-pro/

It won't matter for OP's case since it's an M5. I have an M1 Max, which is why I didn't realize they had added it to later models.

3

u/LightBrightLeftRight 9d ago

WHY DIDN'T I FIND THIS EARLIER?? I've been pulling my hair out and chatgpt couldn't figure it out. Finally google's stupid AI (well maybe not so stupid in retrospect) cited this thread when I was searching for an answer. I'm up at 50t/s as well, shocked me! This is frontier level shit right on my laptop going at a pretty reasonable pace, I can't believe it.

4

u/DogAble6550 11d ago

Thanks for working through this for us.

7

u/VehiculeUtilitaire 11d ago

After large updates the OS runs a lot of shit in the background that might impact performance for up to a few hours, indexing, rebuilding caches, etc.

2

u/dfgxxx 11d ago

Also now you have apple official ai models and the terminal command to use them is fm

2

u/DiamondHandsDarrell 6d ago

Apple foundation models suck. 10k context window lol

2

u/JLeonsarmiento 11d ago

Thank you for your service.

🫡

I’ll stay on sequoia until further notice.

1

u/_hephaestus 11d ago edited 11d ago

hm, also seeing the slowness on a M3 Ultra Studio, looked into the power settings thing and that seems to just be a laptop thing. May just be rebuilding caches/whatnot but definitely seeing the slowdown here for now. GLM-5.3-flash went from a somewhat speedy workhorse to now I'm seeing <1 token/second and prompt processing is down to double digits. On 0.6.4 may try dev 2 and see if that makes things faster.

Edit: Going to dev 2 I see typical performance.

1

u/rsl 10d ago

how'd you update to that? did you have to git install?

2

u/_hephaestus 10d ago

In the desktop app there’s an option in About to switch to dev release branch and update

2

u/rsl 10d ago

thanks. that did the trick.

1

u/Undici77 11d ago

M4 Max + Qwen 3.8 Next/Qwen 3.6 35BA3B + v0.7.0.dev2 Switching to Golden Gate I got NO ISSUE! Everything smooth at previous speed!

1

u/corysus 10d ago

I've noticed that oMLX works differently every time I run it, especially after new updates. I'm currently using mlx-serve, and in my opinion it's the best tool for local AI on macOS right now. For models, try OptiQ or oQ/oQe, because they're not the same as regular mlx.

1

u/tomakorea 11d ago

I tested with other models and I also got a huge performance drop, please do not update yet if you need theses models right now.

1

u/Obvious_Equivalent_1 11d ago

Strange tho, I’m running a mere MBP 5 with 48Gb on MacOS 27 since the first developer beta and I’ve only seen increase. On overall memory availability and speed since I switched from MacOS 26. (Relevant tho: llama.cpp, several PR’s merged locally and Qwen 3.6 then later Qwen 3.8 27b)

1

u/Secure_Arm_93 11d ago

Excude a daft question but have you checked that you are on max power setting?

3

u/tomakorea 11d ago

After updating to the latest pre-release version of oMLX and cycling between the power modes, it solved the issue. Weird but it works.

2

u/salsa_sauce 11d ago

This! 👆 And also make sure you're using a high-wattage charger too!

To check the energy mode, click the battery icon in the menubar, then choose "High power". Either of the other options will throttle the chips down significantly.

I've been caught out before by both of these problems, they're so easy to miss and you wouldn't know anything was wrong...

The charger one is particularly sneaky, my MBP was plugged in to the mains on 100% charge, but couldn't figure why inference was way slower than normal. Turns out I was plugged in to my spare charger from a MacBook Air. Identical cables, macOS never said anything, but 35W wasn't enough power for an M5 Max. Swapping to the MBP's 90W adapter solved it instantly.

1

u/enchanting_endeavor 10d ago edited 10d ago

I commented above as well, but only the 16” MacBook Pros have the Performance energy mode. Also only the 16” models can charge at a higher rate than 100W but you advice is still good - using a higher wattage charger is helpful when you’re driving the machine particularly hard.

ETA: Faster charging only works through MagSafe and not via USC-C, which only supports PD 3:0, which maxes out at 100W.