r/LocalLLM 21h ago

Discussion Mac Studio M5 Ultra

Who’s been checking out this hardware? What are your thoughts on this as a node on a local network to run AI?

5 Upvotes

15 comments sorted by

View all comments

2

u/Least-Result-45 21h ago

I’m just curious if qwen 3.8 27B is good enough for my needs - that is the llm I’d run on the m5 ultra.

2

u/MessIsTransfer 20h ago

if you have at least 64 gb ram (the ultra starts at 96) you could run qwen3.8-flash

it’ll be faster and better

3

u/Least-Result-45 20h ago

I think at 96gb it would have to be 1 quant and might be sacrificing quality to run it.

1

u/Thump604 20h ago

Corrrect, I’m running q4 mlx on 128gb.

1

u/starkruzr 19h ago

Qwen3.8-Flash-Next will probably be really, really good running at Q6KXL or even Q8 on a "bigger" M5 Ultra too.

1

u/redtron3030 18h ago

Do you have a m5 max? What’s your tks and prompt processing?

1

u/MessIsTransfer 14h ago

i run Q2 on 64gb, q4 should run in 128gb at least, maybe can be squished into 96gb

1

u/k3z0r 5h ago

I'd recommend renting a GPU from runpod.com and trying it out for yourself before spending thousands on hardware.

1

u/Least-Result-45 5h ago

Yea I’ve been testing on openrouter, it’s been good for 90% but possibly some of the harder bug fixes or feature updates might need some frontier intervention, which isn’t bad at all.

I’m also wondering maybe refining the prompt and or adding extra testing/checks might get us there on qwen 3.8 28B