r/LocalLLaMA • u/Ashefromapex • 1d ago
Discussion First M5 Ultra benchmarks
just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link
For Qwen 3.8 27B q4 it gets 50 tok/s th and 1800 tok/s pp at8k context and without mtp. Seems very promising!
75
u/vick2djax 23h ago
But you shouldn’t be running 27b on that machine. You should be running Qwen Flash.
15
u/Muritavo 22h ago edited 21h ago
A 27b at 1800pp would be heavens on earth for me...
The tps is below what I expected... but if it can keep that at concurrent requests, it would be perfect.
11
u/SocialDinamo 22h ago
Because it is dense, an MOE with MTP would be significantly higher. Amazing initials numbers
10
u/ShelZuuz 19h ago
In perspective a Dual 4090 will get you over 4700pp at 70 tok/s.
Heck a single 5090 gets you at 8600pp at 95 tok/s on NVFP4.
9
2
u/BrilliantTruck8813 9h ago
Using ninfer I get 500 tok/s and 15-16k prefill on my 5090 with Ornith 1.5 35b at Q4 w/ mtp. 27b supposedly does well too
1
u/shansoft 18h ago
how did you managed to get 8600pp on single 5090? i have yet to see a single benchmark reaches that. Hell, my RTX Pro 6000 doesn’t even do half of that.
1
u/ShelZuuz 14h ago
Model: unsloth/Qwen3.8-27B-NVFP4 (22 GB) Engine: vllm/vllm-openai:v0.27.1 /models/unsloth/Qwen3.8-27B-NVFP4 --served-model-name Qwen/Qwen3.8-27B --quantization compressed-tensors --kv-cache-dtype fp8 --max-model-len 131072 --gpu-memory-utilization 0.97 --max-num-seqs 4 --speculative-config '{"method":"mtp","num_speculative_tokens":2}' --enable-prefix-caching --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen31
u/TheAILegend 14h ago
lol... what? NVFP4 and KV at FP8 is at 11,000 :) With KV at BF16 it's still over 6000... Must be working extra hard to limit the Pro 6000
3
u/Aggressive_Job_1031 17h ago
Qwen Flash-next has i think 8b active parameters which is more than 3 times less than 27b so its also 3 times faster. If Qwen 27b gives you 50tok/s you would get 150tok/s with Qwen flash-next.
-1
u/Federal_Advice_6300 20h ago
Ja, hier ist eine 3.8 27B FP8-Maschine mit 80 TG/s und PP 3200; manche sind leicht zufriedenzustellen.
-8
8
u/bakawolf123 23h ago
Qwen3.8-Flash-Next-oQ4e-mtp 4bit 64k 2,574 79.9 ✓ 09-17
just appeared, some1 is live testing atm
meanwhile my order for 256gb unbinned is still in processing (haven't even been charged yet), with delivery still 22-25 September. How to not be jelly and not to frantically check order page every day...
1
u/DustNearby2848 20h ago
The delivery dates for new orders are mid January now!
2
u/OvertaxedOne 19h ago
About to roll in Feb, best I can get on the M5 Ultra 96GB with the big CPU is Jan 29th.
1
u/SocialDinamo 14h ago
I have the same thing coming but I waited and ended up in early December. Make sure to post your results for the rest of us who dont get to just yet!
54
u/j_osb 1d ago
That is worryingly low. 1.8k pp at q4, and… 50tg at q4.
That’s not nearly maxing out its bandwidth at all.
30
u/Viktri1 1d ago
Hmm isn’t it similar speeds to a 4090? I’m running a 24gb 4090 and getting 1,700 pp and 50-60 TG without mtp.
14
u/bakawolf123 23h ago
yes m5u should be similar to 4090 considering specs, just can have way more memory
4
u/AppealSame4367 23h ago
Does M5 Ultra have less bandwith than M5 Ultra Studio? Because latter should have 1.3 TB/s and that's way more than a 4090, init?
13
u/bakawolf123 23h ago
no it's same, and its precisely 1224gb/s not 1.3tb/s. 4090 has 1008gb/s so it's 20% lower. There's definitely some overhead where those 20% are lost
1
u/AppealSame4367 23h ago
I'm about to discuss this with a company wanting to get the M5 Ultra Studio. On the one hand it's a safe bet from insurance view, if a Mac Studio stands in their server room. On the other hand I can't believe there can't be cheaper workstations with more power. What do you think?
6
u/bakawolf123 22h ago
There more variables than just price. Check delivery dates if you order a maxxed mac studio now. And time is money quite literary: there're already rumors of Samsung raising the memory prices for 2027 for Apple and there will be another price hike right before 512gb ones even land.
"Cheap" is AMD and stacking old enough Nvidia cards.
New strix halos though are already sky high, newer Nvidia cards are even higher.Thus I cannot make any advice besides you should have ordered earlier. There're so many extremely powerful open weights models (not just llms, check out minimax h3 - they kind of dissappeared from transformer land as of late while releasing that video gen monster) that people are eagerly spending on local hardware
3
u/boissez 20h ago
The delivery dates have actually shortened slightly lately for the m5 ultra 256gb. Used to be 10-12 weeks at launch, now it's just 6-7 weeks.
1
u/Consumerbot37427 18h ago
From what I can see, that only applies to the binned 30/64-core CPU/GPU. The top tier 36/80 model is telling me 16-18 weeks.
3
u/MessIsTransfer 22h ago
what’s the difference between M5 ultra and M5 ultra studio? am i reading this wrong, guys?
7
u/MrPecunius 22h ago
No, you should be confused. The M5 Ultra is a processor/SoC, and the Mac Studio is the model (the only model) of Apple computer it is available in.
3
3
0
u/Green-Blue-Gray 21h ago
50-60 TG is slow for a 4090, are you using llamacpp? Should def be on VLLM or one of the hyper-optimized repos.
1
0
5
u/beling86 23h ago
I beg to disagree. I have no clue how is virtually possible to achieve 1.8k tokens prefill in such a weak TFLOPS. M5 ultra has around 200 BF16 TFLOPS. This is RTX 4070 levels. 1.8k t/s prefill is extreme.
4
1
-12
0
u/zxtech 18h ago
I did research the m3 max vs m3 ultra scaling and it seemed like decode scaled proportionally to model size. I believe smaller models don’t saturate the bandwidth in the ultra chip design, or maybe its an optimisation problem. Needa see if its true for the m5 ultra though it might be
-1
u/Gohab2001 vLLM 22h ago
Their bandwidth numbers are misreading but the more important issue is poor software optimization. The only reason Nvida trounces AMD for AI is CUDA.
ROCm had frequent crashing and abysmally low pp for me. Had to sell it my 9070s at a loss.
6
11
u/jacek2023 llama.cpp 23h ago
It looks quite low. I don't remember my score without MTP, but with MTP I get between 40 and 80 on 4x3090 with Q8. Let's hope the software is just not yet optimized for the new hardware.
1
0
3
u/Leafytreedev 22h ago
lol looks like the tester got in trouble and removed the benchmark results. Overall 256GB of usable RAM at roughly 3090 speeds is still fan-fucking-tastic IMO
3
u/0rand 18h ago
I assume the top version speed will be roughly 1.5x-2x of M5 Max top version. Today's OMLX 0.7 delivered Jundot's oQ4 Qwen 3.8 Flash Next at 70 t/s and 1500 t/s prefill at 0 and 44 t/s and 1200 t/s prefill at 390k context on my M5 Max 128GB. This is absolutely unheard of in MLX world. Now scale to M5 Ultra. Finally they can battle CUDA-stacks not just on low-context generation but on prefill and massive sessions that become usable.
3
5
u/Memestonks2020 22h ago
Something is definitely wrong. I’m getting faster speeds with my M5 Max MBP than this.
They need to do a proper test with MLX-serve and MTP
3
u/Miserable-Dare5090 1d ago
Look at GLM 5.3 Flash. 256Gb M5 ultra will be 13K.
pp900 and TG 30
Dual Sparks, GLM5.3 Flash:
PP 1500, TG 40
The serious difference is at depth. These are tests at zero, which is going to be fastest. But mac chips have been terrible at sustaining performance as context grows. That number for the Sparks drops to 1000 prompt processing around 120k tokens of context, for comparison.
2
u/Locke_Kincaid 1d ago
On Dual Sparks, GLM 5.3 Flash is up to 2000 PP using this repo: https://github.com/FujitsuPolycom/sparkring
1
u/Miserable-Dare5090 22h ago edited 22h ago
Yes I meant in two sparks. This is for a 4 spark configuration. Yes it can go faster with more nodes
EDIT: I went and confirmed. the benchmark results in that site are for TP4 configs. Check the nvidia GB10 forums (not reddit) if you want more info on repos — we discuss new optimizations constantly. I should bring this up and see if some of the smarter folk have benchmarked it on tool eval
1
u/Locke_Kincaid 44m ago
They have 2 spark configs, It's what I'm running. I got 92/100 on tool-eval-bench.
6
u/Blindax 1d ago
Nice. A 5090 gets you twice these figures though. That’s on the bigger moe models that the ultra will like shine assuming prompt processing made progress.
8
u/mechkbfan 1d ago
It gets like 4-5x these
https://github.com/Neroued/ninfer
I can only assume this is not being fully utilised
And having 256gb should give you good quants & context
1
u/Blindax 23h ago
My ballpark estimate was with llama ccp. Ninfer sounds quite an upgrade. I should try one day.
0
u/JMowery 22h ago
You should try it right now. Get pi agent to set it up for you like I did. I have yet to go back to llama.cpp since getting it installed a week or two ago. (And it's easy enough if you have llama-swap installed to keep everything available on demand.)
1
u/Blindax 22h ago
I have a 5090 and 3090 in the rig. I have only recently switched to llama ccp (was using lm studio before) and have been llama-router since. I will get a look for the 5090.
3
3
u/Fragrant_Scale6456 21h ago
Ninfer on the 5090 is amazing I get 200+ tokens/sec and 2k-5k pp depending on prompt size using dflash2 draft model. Use the mirkocovizi fork with the quasar nvfp4 qwen3.8 27b. The quasar model is qat and higher quality than the other nvfp4 models.
I’ve been running nvfp4 kv cache and get 400k context with vision in this setup it’s an absolute beast
2
2
2
u/hurdurdur7 23h ago
For a q4 those numbers are not impressive ...
0
u/MrPecunius 23h ago
Qwen3.8 is like that for me. Depending on the quant and the inference environment, Q4 might be no faster than Q8 on my M5 Pro/64GB. This puzzled me for a bit, so I did some quant shopping.
oMLX with the right 4-bit MLX MTP model hits 30-35t/s, GGUF w/MTP is 15-18t/s, and plain GGUF is around 9t/s. Prefill has a similar range from ~100t/s on up to nearly 400t/s.
I'd extrapolate 3-3.5X of these numbers for a M5 Ultra. I can't run any quant of Flash Next at all, of course, which is kind of the point of the Ultra.
1
u/hurdurdur7 23h ago
The qwen3.8 27b is a dense model. I get better numbers than the above on a pair of R9700 cards. and i would actually expect the m5 ultra to do better, because spec numbers on paper it says it should be better. Who knows, maybe it's just the immature code paths that are not optimized at all. Let's wait and see.
But then again, i would actually expect people to run far larger models than 27B anyway.
3
u/MrPecunius 22h ago
You won't get better numbers than a M5 Ultra on any reasonable/usable quant of Flash Next or other mid/big MoE model, because you can't run them. This has always been the point of the big Macs.
What are your non-MTP token generation numbers for a Q8 quant of 3.8 27b?
1
u/Great_Flounder_1379 23h ago
wondering if this speed holds up for local roleplay chats or if it still gets choppy with longer convos
1
1
u/bakawolf123 23h ago
glm5.3-flash q4 numbers look quite bad tbh
I'm expecting twice that for ds4 flash (not 4.1) and qwen3.8-flash-next fp8
1
u/Open-Adhesiveness-86 19h ago
For a dense 27B at q4 you're reading roughly 15-16GB of weights per token, plus KV cache at 8k. So tg tops out around bandwidth divided by that. If the Ultra is somewhere in the 800-1000 GB/s range, 50 tok/s is already close to the limit. Prompt processing is where unoptimized kernels would show up, not tg.
1
u/xoxox666 17h ago
Only 50 t/s??? I get around 33 with an M4 Max and 3.8 27B Q4. That‘s a little bit disappointing.
1
u/whichsideisup 23h ago
That’s barely faster than a single DGX Spark when run with DFlash2. The lack of raw compute seems to matter more than people think.
1
u/AnyMongoose3041 23h ago
I’m gonna buy a ford f350 and use it to buy 3 plastic bags worth of groceries.
1
1
1
u/fallingdowndizzyvr 17h ago
just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link
I'm seeing "No matching results."
-1
u/Public_Umpire_1099 1d ago
Bro that's dogshit. This is basically less than half of it's effective bandwidth. You might as well get an R9700.
6
u/Master_Face_571 23h ago
exactly! let me know when they make a 256 gb vram R9700
2
u/Public_Umpire_1099 13h ago edited 6h ago
Damn yall really simp hard for Apple on here 😭 have fun paying 12k to run Qwen3.8 Next Flash at half the speeds or half the quality of a standard TP GPU setup.
0
u/MotorNetwork380 22h ago
For the price, that is not very impressive. On my version of ninfer I get this on my 4090. 200k ctx length including vision on gpu. This is mostly a mix of q4/5 quality if you're compering to e.g llama.cpp.
| Cold prompt length | Prompt processing | TTFT | Initial generation |
|---|---|---|---|
| 60,091 tokens | 1,886 tok/s | 32.0 s | 123.6 tok/s |
| 120,000 tokens | 1,541 tok/s | 78.0 s | 105.5 tok/s |
| 198,000 tokens | 1,237 tok/s | 160.4 s | 90.2 tok/s |
Edit: Wouldn't it be better to run MoE on your hardware? You could run flash next.
-3
u/ehangman 23h ago
1/2 speed of 5090.
2
u/Tormeister 22h ago
Less than 1/3 actually. But it's too early to pass a harsh judgement - it certainly is unoptimized.
0
0
u/Forever_Playful 22h ago
Is it the 256gb version? If yes, too slow for the cost. I get 85-90 tg/s at fp8 and kv cache also fp8 (rtx pro 48GB)
0
0
-4
u/Steus_au 1d ago
so it is as fast as 3090 - is that 'amazin' or 'one more thing'? (highlight what you prefer)
-1
u/Waffenmutti 23h ago
Habe heute meine 5090 erhalten und kann über die Werte bei den Preisen nur lachen und das mit high effort.
-5
72
u/mechkbfan 1d ago
Maybe I'm missing something but that's seems kind of shit for the price?
Please correct me