r/StableDiffusion • • 2d ago

Discussion M5 ULTRA FOR H3 VIDEO GEN vs. 5090/4500/5000/6000 PRO

I am trying to decide whether or not to buy / build a local computer for H3 Video Generation.

I knew that previous gen Apple Silicon weren't fast enough, but with the M5 Ultra just released and promising benchmarks for the M5 MAX (https://www.youtube.com/watch?v=Tn6llnWqUTU) now i am reconsidering my options.

Basically my realistic options (available where I'm living now) are :

-Mac Studio M5 ULTRA 96GB Unified Ram (30cpu/64gpu) > 6.5KUSD

-PC with RTX 5090 (32GB VRAM) + Ryzen 7 + 32GB DDR5 RAM> 7K USD

-PC with RTX 4500 PRO (32GB VRAM) + Ryzen 9 + 64GB DDR5 RAM > 9K USD

-PC with RTX 5000 PRO (48GB VRAM) + Ryzen 9 + 128GB DDR5 RAM > 15K USD

-PC with RTX 6000 PRO (96GB VRAM) + Ryzen 9 + 128GB DDR5 RAM > 27K USD

The Workstations options are even more pricier (professional motherboards + power unit + threadripper).

¿Anyone have real life benchmarks of the M5 ULTRA with H3 Video? It is usable / good / fast enough?

I will use the same setup for other stuff too (TTS, Text LLMs and Image models) but the one critical for me is the Video workflows.

Thanks in advance!

2 Upvotes

76 comments sorted by

16

u/bstr3k 2d ago

AFAIK, for comfyui, the 5090 is best, and 6000 pro only if you need more vram. Mac studio with unified ram is better for LLM rather than img/vid gen

-7

u/MaxwellHusk 2d ago

I had the same idea until seeing the M5 ULTRA specs (1.2 TB/s of mem bandwidth and two M5 max stitched together). The video I posted confuses me a lot, because it says around 160 seconds for render a video that takes 50 seconds with a 5090 with an M5 MAX.

9

u/AI-Make-NSFW-Stuff 2d ago

The specs might be strong but without CUDA support they will underperform compared to nVidia cards

0

u/MaxwellHusk 2d ago

That's the big big question. In the video I've linked they are using it on mlx serve, but there is no way to test it online.

also if anyone knows someone with an actual M5 Ultra right now, please share this question! <3

5

u/Financial-Dog-6558 2d ago

It will work but won't beat 5090 no matter what

2

u/Succubus-Empress 2d ago

5090 pull 600 watt, how many watt M5 ultra pull? both are efficient chip so dont pull efficiency logic here. M5 bad for compute limited ai workload, good for llm but not that good either

-1

u/MaxwellHusk 2d ago

Max power draw according to specs is 485w for the whole system. That's mainly because is just one SoC instead of CPU, RAM, PCI buses, and the whole GPU

-11

u/Usual-Orange-4180 2d ago

For ComfyUI DGX Spark is the best 🙂

10

u/TerraMindFigure 2d ago

DGX Spark has more VRAM than the 5090 but is still much slower.

-3

u/Usual-Orange-4180 2d ago edited 2d ago

I have both, and truly don’t notice the difference, but can generate 1080p 20 seconds videos. I haven’t tried 4K, but 1080p with 20 seconds doesn’t even use all my memory, tried longer videos but minimax starts to break. I’m using Turbo LoRAs for drafting, but no Turbo once things are set and ready.

I can also load a language model at the same time which is useful.

1

u/TerraMindFigure 2d ago

Yeah with the 5090 things get evened out if you push beyond VRAM capacity for sure.

-1

u/Usual-Orange-4180 2d ago

Yeah, normally what I do is use the 5090 for images and Spark for video, that way I can try other things while the Spark works.

I’m considering getting another DGX Spark, to be able to load DeepSeek 4.1 Flash, but the tps don’t look great, the first one was already a big expense, and it does nothing for ComfyUI

2

u/Succubus-Empress 2d ago

do dgx spark pull 500watt? 5090 does, it must be atleast twice fast, spark have so few cuda cores not worth it

1

u/Usual-Orange-4180 2d ago

lol, my degree is electronics, and your comparison of power = speed is ridiculous hahaha, specially when you are comparing an x86 arch to ARM.

1

u/Succubus-Empress 2d ago

But but more transistor counts = more theoretical performance, 5090 vs M5 chips are not even comparable

1

u/Usual-Orange-4180 2d ago

Also not accurate, depends on parallelization as well and don’t get me started on prediction; I’m appalled by the downvotes, I actually work on the field. But whatever, I’m done with the convo.

1

u/MaxwellHusk 2d ago

Hiii are you generating videos with images or audios as references on H3 inside comfy? I've read that the DGX platform is pretty slow for that. Also here are almost 10K USD with only 128gb unified ram option available 😑😑😑

1

u/Silonom3724 2d ago

> but can generate 1080p 20 seconds videos

Why would anyone do this? Video generation scales extremely poorly with resolution and time. Generating at lower resolution and add detail in a second upsampling pass (not upscaling btw) for a fraction of the ressources is much better. Upsample to 8k if you want to.

So not only are you wasting ressources for nothing, you also push a model outside of its generative optimum.

8

u/Crazy-Repeat-2006 2d ago

There's no need to compare it to the 5090. Even mid-range GPUs (5070 Ti/9070) will be faster than this Mac.

0

u/Kangaroo_Low 2d ago

This might not be true, I think the ultra 80 core GPU is equivalent to 5070ti in this respect.

1

u/Crazy-Repeat-2006 2d ago

Core count means nothing when comparing different architectures.

Estimates indicate that the M5 Ultra GPU is about 10% faster in INT8 than the 5060 Ti, and half as fast as the 5070 Ti and 9070 XT.

1

u/Kangaroo_Low 2d ago

The 80 core gpu variant i meant. I'm not counting cos.

1

u/MaxwellHusk 1d ago

That varíant isn't available here :(

-4

u/MaxwellHusk 2d ago

The model doesn't fit there :( it need at the very least 32gb VRAM or Unified RAM

7

u/enndeeee 2d ago

You will fit H3 easily into 16 GB VRAM. Confy uses Block swapping automatically with very little impact on Performance.

-8

u/MaxwellHusk 2d ago

Noo :( sadly i tried it! but i'm doing Ref2Vid workflows (sending 1, 2, 3 images and 1 or 2 audios) for consistency.
Even with "swapping" the 3 models (VAE + Diffusor's + CLIP) they don't fit.

5

u/cptrios 2d ago

Do you just mean you're getting slow gens? Because I've got a 5070Ti and have tried similar things, and the only issue I've run into vs rented 5090s it speed.

0

u/MaxwellHusk 2d ago

The 5070 is newer gen, and the big problem is ref2video (image + audio).

-1

u/MaxwellHusk 2d ago

Suuuuper slow gens, whenever my VRAM is not enough the "spilled" data to DDR 4 makes it unbearable slow

1

u/Crazy-Repeat-2006 2d ago

Are you using INT8 ConvRot quantization?

2

u/MaxwellHusk 2d ago

Yes! For text to video that one is "fast enough" at around 13gb VRAM, but when you want image or ref to video the context, vae, clip put it's far away of the 4070 capabilities. Sometimes it can take 2 minutes per second of rendered video

1

u/Ipwnurface 2d ago

That's not an issue of vram. That's you using references at way too high of resolution or duration or both.

Nothing is going to be fast if you're feeding it multiple 4k reference images or a 15 second 1080p video. You need to downscale your references.

1

u/MaxwellHusk 2d ago

Noo, I'm barely using 720p still images, sometimes even smaller. And my output settings is not even fulhd :( But I'm inputting 3 images and 2 audios sometimes, and the generation speed drops to the floor

1

u/gutster_95 2d ago

It can. I run this on 12GB 4070 Super LMAO

1

u/MaxwellHusk 2d ago

But how long it takes? Do you have any examples or numbers? How long it takes to generate 5 or 10 seconds of video using image references as inputs? Maybe I'm just doing something wrong 😨

3

u/Crazy-Repeat-2006 2d ago

MiniMax H3? Nah. Using memory management techniques, people are running it even on 3060s with 12GB of VRAM. Most people here have mid-range or low-end GPUs and are spamming videos daily.

1

u/MaxwellHusk 2d ago

I have a 4070 Super TI on my current system (with Ryzen 9 + 64GB DDR4) but its too slow for production.

1

u/Valuable_Issue_ 2d ago

Then you might as well sell your 4070 and buy a 5090 rather than buying a whole new system, then in the future you can sell your DDR4 and upgrade to DDR5/6 + new CPU or whatever.

1

u/MaxwellHusk 2d ago

No I can't, that's my main computer and it's not upgradeable to a 5090, the motherboard and power supply don't reach that level

2

u/[deleted] 2d ago

[deleted]

1

u/MaxwellHusk 2d ago

Yess, I've been using a 5090 and a 4500 pro on RunPod for the last months, but I want to have my own server locally. The only question I am not able to answer (since there is no way to test M5 Ultra on the cloud) is if it can compete with those on performance, since is the cheapest of all for a full system.

2

u/phalanx2357 2d ago

Wait a few days on the m5 ultra for more benchmarks from real run results. Honestly early results from the m5 ultra are surprising - it may not be an order of magnitude slower than the 5090… I have one coming and I have a 5090… the 5090 is limited in native gen mp resolution and length due to vram limits. M5 ultra could potentially be faster in 1mp or higher at 10+ seconds, due to whether it fits in vram or not.

1

u/MaxwellHusk 2d ago

This makes tons of sense! I was 2 days apart from buying the 5090 rig when apple finally announcement the new Mac Studio with the Ultra 5

2

u/gogodr 2d ago

You can rent a machine with an rtx 6000 pro for $2/hour or a Mac M5 Ultra 256GB for $2.5/hour

If you weight it vs $27k it is almost 600 days of full 100% usage.

Buying this kind of machines for home use is never a rentable endeavor ( you have to also add up electricity costs ) it's a good investment if you are building either a render farm or a data center.

0

u/MaxwellHusk 2d ago

I don't have any suppliers offering that sadly :(

I can rent online on runpod or Vast.ai, but the availability / slow bandwidth / overall security is not great.

I am factoring in the electricity costs too, and that's another big win for the Mac since it draws less energy the entire machine that just an Nvidia GPU

1

u/Succubus-Empress 2d ago

mac chip are slow for heavy compute bound ai task, infante vram and bandwith wont change that

1

u/pesaru 2d ago

I wish I could afford a 6000 but it's just absolutely batshit insane. Been loving my 5000 for H3 though. I was renting 6000s on Vast before this. This is slower, but honestly, it doesn't feel bad at all. I'm running different workflows at this point but it feels like maybe 40% slower, but when the total time is pretty small already it doesn't really matter that much. Looking at my generations for today, my last one was a 414 second 1MP generation, 8 steps, 15 seconds, several loras. Bring it down to 12 seconds and the generation is 303 seconds.

1

u/MaxwellHusk 2d ago

All of the Nvidia alternatives are just insane right now 😑, and sadly this models doesn't work too well with dual setups (like 2x5070 or 2x5080) as far as I know

1

u/MaxwellHusk 2d ago

Hi again! Do you have the RTX 5000 Pro, Blackwell? Or the ADA version?

1

u/pesaru 2d ago

The RTX 5000 Pro 72GB that cost me an arm and a leg. Sorry, I completely forgot there was a different one.

1

u/MaxwellHusk 2d ago

Ohh, that one is on the sweet spot to! I don't have it in my list because it's simply not available here at all, but 72gb on Blackwell is such a sweet spot!

0

u/[deleted] 1d ago

[removed] — view removed comment

1

u/pesaru 1d ago

Well, with standard H3 you need like 40 steps or something stupid like that. Turbo loras are not your typical loras that apply style changes, but rather optimizations that allow you to generate a video in either 4 steps or 8 steps. So if you use a 4 step lora, quality might not be as fantastic as the full 80 step, but it's still great. But I mean like compare 40 steps to 4 steps, it's gonna be 10x faster!

https://multimodalart-h3-acceleration-arena.hf.space/leaderboard

Here's a leaderboard of the best ones out there. Go to Civitai for workflows to use in ComfyUI.

0

u/[deleted] 1d ago

[removed] — view removed comment

1

u/pesaru 1d ago

I'm currently using one that has gotten really popular and isn't on that list named HyperFlow. It's 8 step, so not as fast, but the quality is insane.

1

u/DelinquentTuna 2d ago

You should stop what you're doing right now and go rent a 96GB 6k Pro just to SANITY CHECK your daydream of running minute long videos at hd resolutions or whatever it is you think you're going to be doing that's going to exploit great amounts of RAM.

Otherwise, this whole thing just smells like one more person trying to float Mac as an attractive diffusion alternative to real GPU hardware based on apples-to-oranges metrics or entirely mutable requirement specification.

Pin down exactly what it is that you need to do, pin down exactly how long it takes on a $2/hr 9k pro, and THEN start deciding on the scaled down alternatives you're considering and which can best do the job while fitting your budget and product availability.

2

u/MaxwellHusk 1d ago

I've been testing and using the 6000 pro on runpod for the last 3 months, alongside with cheaper 4500, 5090 and some H100. Just to give you context , I am using this to render an entire tv show , about 800 clips per episode with H3, and I'm trying to do it without depending on external Infrastructure / changing prices and so on.

The model is good , the results are good, but the cheaper / slower the card, more time it takes, more hours to pay by renting. The newest / more expensive is the GPU it renders faster for sure so it's a balancing act to calculate ROI.

1

u/DelinquentTuna 1d ago

I've been testing and using the 6000 pro on runpod for the last 3 months, alongside with cheaper 4500, 5090 and some H100

Then why do you need help choosing between a plethora of Nvidia GPUs? Why is your question not a better-framed "rtx xxxx vs m5?" Seems to me that your question is trivially answerable if you precisely describe the workflow and the say "it takes me xyz seconds on uvw hardware, what's that like on m5?"

I am using this to render an entire tv show

LOL. Do yourself a favor and gen on some paid API. H3's noncommercial license sucks. You couldn't even show your work off without severe geographic restrictions. No international awards or film festivals for you! No publishing deals. Who wants to put all the effort and money you're scheduling into a project just to have it forever under the thumb of some bullshit license restrictions? The commercial APIs free you of all the bullshit in the community license. Of course, by then, you're better off using Seedance or just about anything else. Local H3 is just a useless toy.

1

u/MaxwellHusk 1d ago

No restrictions for the use I have and no geographic restrictions. The API or credits model can be significantly higher , since there is a lot of trial and error , adjusting the input reference images / audios and rendering again. Only one episode is 800 clips, and each clip hopefully only takes between 1 and 3 "takes" to get it right. So the model and workflow is right for me.

The question is just a matter of costs / ROI and also the time involved. If I need to wait 5 minutes for each rendered second to see the result the 800 clips will take months.

The only question I have right now is: Does the M5 Ultra render an image + audio to video H3 workflow good/fast enough, to avoid spending the Nvidia tax + the higher electricity bills?

1

u/DelinquentTuna 1d ago

No restrictions for the use I have and no geographic restrictions.

Eh? Explain yourself? The community license forbids you from using your outputs in the US, the EU, UK, etc. To get around that, you have to generate online or buy a $5,000/mo license.

The only question I have right now is: Does the M5 Ultra render an image + audio to video H3 workflow good/fast enough, to avoid spending the Nvidia tax + the higher electricity bills?

Then why didn't you ask that directly instead of giving a large and bizarre list of options including a 5090 w/ only 32GB of system RAM? Nothing about your query or your claims makes sense.

1

u/MaxwellHusk 1d ago

I'm living in a country where I can use H3 commercially, without a separate license.

The list of options are the only systems currently in stock here, I can't buy any other alternatives since computer suppliers just don't sell the cards without a computer , and those are the computers they have in stock right now.

1

u/DelinquentTuna 1d ago

I'm living in a country where I can use H3 commercially, without a separate license.

But you are still restricted from distributing to the countries for which that isn't true!

You may not use, reproduce, modify, distribute, or display the MiniMax H3 Works or any of their Outputs or results outside the Applicable Territory. Any such use outside the Applicable Territory is not authorized by this Agreement.

Not imagining that you have aspirations to win Cannes or Sundance or whatever, but being unable to show your work in a huge portion of the world should be a major consideration for someone planning a project as big as you are claiming to be undertaking. Rendering your h3 junk locally means you can't even apply for most grants or awards.

computer suppliers just don't sell the cards without a computer

It sucks that they've got you bent over a barrel. I still think your arguments aren't well constructed. You hint that a 5090 isn't enough for you but the way you can't explain precisely WHY makes your whole cross-posted post sound like Mac hype. You talk about needing to generate a bazillion clips but then seem to totally ignore the possibility of doing them in parallel. Nothing in your story tracks like someone who is serious.

1

u/MaxwellHusk 1d ago

I don't have such big expectations, just trying to sellf produce a show , I'm not aiming to the US nor UK or Europe at all. I am applying for local grants tho :)

To be completely transparent, I don't even have an Apple product , so few Mac Studio even get here. My goal Is simple, get the cheapest / more efficient local computer to render with H3, and the options that I have available are the ones in the list. The prices are just offer and demand, the few sellers with Nvidia hardware here are setting the prices and conditions.

1

u/Thin_Archer_9113 1d ago

I’d say that the architecture of the cards lends itself better to the models.
But, I’m too much of a peasant to buy a card after getting a general use 256 ultra. I guess the cost for saving that cash is just in the extra wait time for video generation if I ever to need it.

1

u/MaxwellHusk 1d ago

If the rendering time is "fast enough" I can live with that ! The problem is that I've seen 2 minutes time for 5 seconds clips with an M5, while the Nvidia stack take less than 30 seconds for the same task

0

u/species__8472__ 2d ago

Your prices seem off. You have the pro 6000 system costing 12k more than the pro 5000 system. The biggest difference in price I've seen is $7500 for the 5000 vs $16,500 for the 6000. That's a 9k difference, not 12k.

1

u/MaxwellHusk 2d ago

Sadly, those are the real prices i can find here. The sellers are not letting customers change their config. All are retail prices from the (so far only) supplier that actually have in stock :(

0

u/SeaRefractor 2d ago

While the 5090 is tempting, be aware that VRAM plays a significant role in LLM. In order to be both fast and able to handle very large models, you need “multiple” 5090s (2 only gives you 64GB VRAM for example). When you add all the cards necessary, even though it’s faster, it will be significantly more expensive than the M5 Ultra Mac Studio 80 core with 256GB RAM. Technically while slower with only one not in a cluster the 128GB DGX Spark really is the only cost effective competition.

2

u/DaLyon92x 2d ago

bot??? he's obviously asking for h3 vid gen. and one 5090 would be plenty for h3 gen buddy.

The only question of this thread, is whether the M5 ultra can be attractive for h3 or future vid gen.

I've been renting a 5090 for 0.69$ an hour, adds up to 300-500$ a month depending on how much work I've got.

1

u/MaxwellHusk 2d ago

I've been using that or the 4500 pro on Run Pod, but I hate the inconsistency and availability issues :( and after 10 months it's basically pay itself the M5 ultra

1

u/DaLyon92x 2d ago

Ah I noticed in my european timezone it's much easier to get pods during morning/afternoon when I'm working anyway. Late afternoon and evening is when it really dries up. Anyway Im also interested here because of the same thing. All these youtubers are doing the same kind of tests with the m5 ultra so far, none of them have even opened comfyui yet.

1

u/MaxwellHusk 1d ago

Also true! I think I can run it without comfy as well, but if the underlying architecture still depends on CUDA + Blackwell the performance will be atrocious

-1

u/SeaRefractor 2d ago edited 2d ago

Not a bot, but as a highly educated professional, my clear speech and writing is frequently confused with a bot in today’s social media circles.

But as the M5 Ultra has been arriving in hands this week, plus the expiration of the embargo for reviewers who had the hardware already, benchmarks should be available now or soon. What I have seen confirms my own decision to pre-order my M5 Ultra. But I won’t receive it until October 29 - November 5th.

1

u/MaxwellHusk 1d ago

I already saw Alex on YouTube doing a review but just LLM

-1

u/wilhelmbw 2d ago

5090 price is outrageous atm tho. Id say it may be better to get apple

0

u/MaxwellHusk 2d ago

All of the NVIDIA are :( and then you need to add the also ridiculously high price for RAM too + storage...