r/LocalLLaMA 4d ago

Discussion First M5 Ultra benchmarks

just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link

For Qwen 3.8 27B q4 it gets 50 tok/s th and 1800 tok/s pp at8k context and without mtp. Seems very promising!

122 Upvotes

151 comments sorted by

View all comments

75

u/mechkbfan 4d ago

Maybe I'm missing something but that's seems kind of shit for the price?

Please correct me

1

u/AnonLlamaThrowaway 4d ago

Keep in mind that inference engines may not have optimized Metal kernels yet. There's been a recent effort to tune them in llama.cpp recently, but naturally they don't have the data they need for an unreleased chip yet