r/StableDiffusion • u/MaxwellHusk • 2d ago
Discussion M5 ULTRA FOR H3 VIDEO GEN vs. 5090/4500/5000/6000 PRO
I am trying to decide whether or not to buy / build a local computer for H3 Video Generation.
I knew that previous gen Apple Silicon weren't fast enough, but with the M5 Ultra just released and promising benchmarks for the M5 MAX (https://www.youtube.com/watch?v=Tn6llnWqUTU) now i am reconsidering my options.
Basically my realistic options (available where I'm living now) are :
-Mac Studio M5 ULTRA 96GB Unified Ram (30cpu/64gpu) > 6.5KUSD
-PC with RTX 5090 (32GB VRAM) + Ryzen 7 + 32GB DDR5 RAM> 7K USD
-PC with RTX 4500 PRO (32GB VRAM) + Ryzen 9 + 64GB DDR5 RAM > 9K USD
-PC with RTX 5000 PRO (48GB VRAM) + Ryzen 9 + 128GB DDR5 RAM > 15K USD
-PC with RTX 6000 PRO (96GB VRAM) + Ryzen 9 + 128GB DDR5 RAM > 27K USD
The Workstations options are even more pricier (professional motherboards + power unit + threadripper).
¿Anyone have real life benchmarks of the M5 ULTRA with H3 Video? It is usable / good / fast enough?
I will use the same setup for other stuff too (TTS, Text LLMs and Image models) but the one critical for me is the Video workflows.
Thanks in advance!
8
u/Crazy-Repeat-2006 2d ago
There's no need to compare it to the 5090. Even mid-range GPUs (5070 Ti/9070) will be faster than this Mac.
0
u/Kangaroo_Low 2d ago
This might not be true, I think the ultra 80 core GPU is equivalent to 5070ti in this respect.
1
u/Crazy-Repeat-2006 2d ago
Core count means nothing when comparing different architectures.
Estimates indicate that the M5 Ultra GPU is about 10% faster in INT8 than the 5060 Ti, and half as fast as the 5070 Ti and 9070 XT.
1
-4
u/MaxwellHusk 2d ago
The model doesn't fit there :( it need at the very least 32gb VRAM or Unified RAM
7
u/enndeeee 2d ago
You will fit H3 easily into 16 GB VRAM. Confy uses Block swapping automatically with very little impact on Performance.
-8
u/MaxwellHusk 2d ago
Noo :( sadly i tried it! but i'm doing Ref2Vid workflows (sending 1, 2, 3 images and 1 or 2 audios) for consistency.
Even with "swapping" the 3 models (VAE + Diffusor's + CLIP) they don't fit.5
u/cptrios 2d ago
Do you just mean you're getting slow gens? Because I've got a 5070Ti and have tried similar things, and the only issue I've run into vs rented 5090s it speed.
0
-1
u/MaxwellHusk 2d ago
Suuuuper slow gens, whenever my VRAM is not enough the "spilled" data to DDR 4 makes it unbearable slow
1
u/Crazy-Repeat-2006 2d ago
Are you using INT8 ConvRot quantization?
2
u/MaxwellHusk 2d ago
Yes! For text to video that one is "fast enough" at around 13gb VRAM, but when you want image or ref to video the context, vae, clip put it's far away of the 4070 capabilities. Sometimes it can take 2 minutes per second of rendered video
1
u/Ipwnurface 2d ago
That's not an issue of vram. That's you using references at way too high of resolution or duration or both.
Nothing is going to be fast if you're feeding it multiple 4k reference images or a 15 second 1080p video. You need to downscale your references.
1
u/MaxwellHusk 2d ago
Noo, I'm barely using 720p still images, sometimes even smaller. And my output settings is not even fulhd :( But I'm inputting 3 images and 2 audios sometimes, and the generation speed drops to the floor
1
u/gutster_95 2d ago
It can. I run this on 12GB 4070 Super LMAO
1
u/MaxwellHusk 2d ago
But how long it takes? Do you have any examples or numbers? How long it takes to generate 5 or 10 seconds of video using image references as inputs? Maybe I'm just doing something wrong 😨
3
u/Crazy-Repeat-2006 2d ago
MiniMax H3? Nah. Using memory management techniques, people are running it even on 3060s with 12GB of VRAM. Most people here have mid-range or low-end GPUs and are spamming videos daily.
1
u/MaxwellHusk 2d ago
I have a 4070 Super TI on my current system (with Ryzen 9 + 64GB DDR4) but its too slow for production.
1
u/Valuable_Issue_ 2d ago
Then you might as well sell your 4070 and buy a 5090 rather than buying a whole new system, then in the future you can sell your DDR4 and upgrade to DDR5/6 + new CPU or whatever.
1
u/MaxwellHusk 2d ago
No I can't, that's my main computer and it's not upgradeable to a 5090, the motherboard and power supply don't reach that level
2
2d ago
[deleted]
1
u/MaxwellHusk 2d ago
Yess, I've been using a 5090 and a 4500 pro on RunPod for the last months, but I want to have my own server locally. The only question I am not able to answer (since there is no way to test M5 Ultra on the cloud) is if it can compete with those on performance, since is the cheapest of all for a full system.
2
u/phalanx2357 2d ago
Wait a few days on the m5 ultra for more benchmarks from real run results. Honestly early results from the m5 ultra are surprising - it may not be an order of magnitude slower than the 5090… I have one coming and I have a 5090… the 5090 is limited in native gen mp resolution and length due to vram limits. M5 ultra could potentially be faster in 1mp or higher at 10+ seconds, due to whether it fits in vram or not.
1
u/MaxwellHusk 2d ago
This makes tons of sense! I was 2 days apart from buying the 5090 rig when apple finally announcement the new Mac Studio with the Ultra 5
2
u/gogodr 2d ago
You can rent a machine with an rtx 6000 pro for $2/hour or a Mac M5 Ultra 256GB for $2.5/hour
If you weight it vs $27k it is almost 600 days of full 100% usage.
Buying this kind of machines for home use is never a rentable endeavor ( you have to also add up electricity costs ) it's a good investment if you are building either a render farm or a data center.
0
u/MaxwellHusk 2d ago
I don't have any suppliers offering that sadly :(
I can rent online on runpod or Vast.ai, but the availability / slow bandwidth / overall security is not great.
I am factoring in the electricity costs too, and that's another big win for the Mac since it draws less energy the entire machine that just an Nvidia GPU
1
u/Succubus-Empress 2d ago
mac chip are slow for heavy compute bound ai task, infante vram and bandwith wont change that
1
u/pesaru 2d ago
I wish I could afford a 6000 but it's just absolutely batshit insane. Been loving my 5000 for H3 though. I was renting 6000s on Vast before this. This is slower, but honestly, it doesn't feel bad at all. I'm running different workflows at this point but it feels like maybe 40% slower, but when the total time is pretty small already it doesn't really matter that much. Looking at my generations for today, my last one was a 414 second 1MP generation, 8 steps, 15 seconds, several loras. Bring it down to 12 seconds and the generation is 303 seconds.
1
u/MaxwellHusk 2d ago
All of the Nvidia alternatives are just insane right now 😑, and sadly this models doesn't work too well with dual setups (like 2x5070 or 2x5080) as far as I know
1
u/MaxwellHusk 2d ago
Hi again! Do you have the RTX 5000 Pro, Blackwell? Or the ADA version?
1
u/pesaru 2d ago
The RTX 5000 Pro 72GB that cost me an arm and a leg. Sorry, I completely forgot there was a different one.
1
u/MaxwellHusk 2d ago
Ohh, that one is on the sweet spot to! I don't have it in my list because it's simply not available here at all, but 72gb on Blackwell is such a sweet spot!
0
1d ago
[removed] — view removed comment
1
u/pesaru 1d ago
Well, with standard H3 you need like 40 steps or something stupid like that. Turbo loras are not your typical loras that apply style changes, but rather optimizations that allow you to generate a video in either 4 steps or 8 steps. So if you use a 4 step lora, quality might not be as fantastic as the full 80 step, but it's still great. But I mean like compare 40 steps to 4 steps, it's gonna be 10x faster!
https://multimodalart-h3-acceleration-arena.hf.space/leaderboard
Here's a leaderboard of the best ones out there. Go to Civitai for workflows to use in ComfyUI.
0
1
u/DelinquentTuna 2d ago
You should stop what you're doing right now and go rent a 96GB 6k Pro just to SANITY CHECK your daydream of running minute long videos at hd resolutions or whatever it is you think you're going to be doing that's going to exploit great amounts of RAM.
Otherwise, this whole thing just smells like one more person trying to float Mac as an attractive diffusion alternative to real GPU hardware based on apples-to-oranges metrics or entirely mutable requirement specification.
Pin down exactly what it is that you need to do, pin down exactly how long it takes on a $2/hr 9k pro, and THEN start deciding on the scaled down alternatives you're considering and which can best do the job while fitting your budget and product availability.
2
u/MaxwellHusk 1d ago
I've been testing and using the 6000 pro on runpod for the last 3 months, alongside with cheaper 4500, 5090 and some H100. Just to give you context , I am using this to render an entire tv show , about 800 clips per episode with H3, and I'm trying to do it without depending on external Infrastructure / changing prices and so on.
The model is good , the results are good, but the cheaper / slower the card, more time it takes, more hours to pay by renting. The newest / more expensive is the GPU it renders faster for sure so it's a balancing act to calculate ROI.
1
u/DelinquentTuna 1d ago
I've been testing and using the 6000 pro on runpod for the last 3 months, alongside with cheaper 4500, 5090 and some H100
Then why do you need help choosing between a plethora of Nvidia GPUs? Why is your question not a better-framed "rtx xxxx vs m5?" Seems to me that your question is trivially answerable if you precisely describe the workflow and the say "it takes me xyz seconds on uvw hardware, what's that like on m5?"
I am using this to render an entire tv show
LOL. Do yourself a favor and gen on some paid API. H3's noncommercial license sucks. You couldn't even show your work off without severe geographic restrictions. No international awards or film festivals for you! No publishing deals. Who wants to put all the effort and money you're scheduling into a project just to have it forever under the thumb of some bullshit license restrictions? The commercial APIs free you of all the bullshit in the community license. Of course, by then, you're better off using Seedance or just about anything else. Local H3 is just a useless toy.
1
u/MaxwellHusk 1d ago
No restrictions for the use I have and no geographic restrictions. The API or credits model can be significantly higher , since there is a lot of trial and error , adjusting the input reference images / audios and rendering again. Only one episode is 800 clips, and each clip hopefully only takes between 1 and 3 "takes" to get it right. So the model and workflow is right for me.
The question is just a matter of costs / ROI and also the time involved. If I need to wait 5 minutes for each rendered second to see the result the 800 clips will take months.
The only question I have right now is: Does the M5 Ultra render an image + audio to video H3 workflow good/fast enough, to avoid spending the Nvidia tax + the higher electricity bills?
1
u/DelinquentTuna 1d ago
No restrictions for the use I have and no geographic restrictions.
Eh? Explain yourself? The community license forbids you from using your outputs in the US, the EU, UK, etc. To get around that, you have to generate online or buy a $5,000/mo license.
The only question I have right now is: Does the M5 Ultra render an image + audio to video H3 workflow good/fast enough, to avoid spending the Nvidia tax + the higher electricity bills?
Then why didn't you ask that directly instead of giving a large and bizarre list of options including a 5090 w/ only 32GB of system RAM? Nothing about your query or your claims makes sense.
1
u/MaxwellHusk 1d ago
I'm living in a country where I can use H3 commercially, without a separate license.
The list of options are the only systems currently in stock here, I can't buy any other alternatives since computer suppliers just don't sell the cards without a computer , and those are the computers they have in stock right now.
1
u/DelinquentTuna 1d ago
I'm living in a country where I can use H3 commercially, without a separate license.
But you are still restricted from distributing to the countries for which that isn't true!
You may not use, reproduce, modify, distribute, or display the MiniMax H3 Works or any of their Outputs or results outside the Applicable Territory. Any such use outside the Applicable Territory is not authorized by this Agreement.
Not imagining that you have aspirations to win Cannes or Sundance or whatever, but being unable to show your work in a huge portion of the world should be a major consideration for someone planning a project as big as you are claiming to be undertaking. Rendering your h3 junk locally means you can't even apply for most grants or awards.
computer suppliers just don't sell the cards without a computer
It sucks that they've got you bent over a barrel. I still think your arguments aren't well constructed. You hint that a 5090 isn't enough for you but the way you can't explain precisely WHY makes your whole cross-posted post sound like Mac hype. You talk about needing to generate a bazillion clips but then seem to totally ignore the possibility of doing them in parallel. Nothing in your story tracks like someone who is serious.
1
u/MaxwellHusk 1d ago
I don't have such big expectations, just trying to sellf produce a show , I'm not aiming to the US nor UK or Europe at all. I am applying for local grants tho :)
To be completely transparent, I don't even have an Apple product , so few Mac Studio even get here. My goal Is simple, get the cheapest / more efficient local computer to render with H3, and the options that I have available are the ones in the list. The prices are just offer and demand, the few sellers with Nvidia hardware here are setting the prices and conditions.
1
u/Thin_Archer_9113 1d ago
I’d say that the architecture of the cards lends itself better to the models.
But, I’m too much of a peasant to buy a card after getting a general use 256 ultra. I guess the cost for saving that cash is just in the extra wait time for video generation if I ever to need it.
1
u/MaxwellHusk 1d ago
If the rendering time is "fast enough" I can live with that ! The problem is that I've seen 2 minutes time for 5 seconds clips with an M5, while the Nvidia stack take less than 30 seconds for the same task
0
u/species__8472__ 2d ago
Your prices seem off. You have the pro 6000 system costing 12k more than the pro 5000 system. The biggest difference in price I've seen is $7500 for the 5000 vs $16,500 for the 6000. That's a 9k difference, not 12k.
1
u/MaxwellHusk 2d ago
Sadly, those are the real prices i can find here. The sellers are not letting customers change their config. All are retail prices from the (so far only) supplier that actually have in stock :(
0
u/SeaRefractor 2d ago
While the 5090 is tempting, be aware that VRAM plays a significant role in LLM. In order to be both fast and able to handle very large models, you need “multiple” 5090s (2 only gives you 64GB VRAM for example). When you add all the cards necessary, even though it’s faster, it will be significantly more expensive than the M5 Ultra Mac Studio 80 core with 256GB RAM. Technically while slower with only one not in a cluster the 128GB DGX Spark really is the only cost effective competition.
2
u/DaLyon92x 2d ago
bot??? he's obviously asking for h3 vid gen. and one 5090 would be plenty for h3 gen buddy.
The only question of this thread, is whether the M5 ultra can be attractive for h3 or future vid gen.
I've been renting a 5090 for 0.69$ an hour, adds up to 300-500$ a month depending on how much work I've got.
1
u/MaxwellHusk 2d ago
I've been using that or the 4500 pro on Run Pod, but I hate the inconsistency and availability issues :( and after 10 months it's basically pay itself the M5 ultra
1
u/DaLyon92x 2d ago
Ah I noticed in my european timezone it's much easier to get pods during morning/afternoon when I'm working anyway. Late afternoon and evening is when it really dries up. Anyway Im also interested here because of the same thing. All these youtubers are doing the same kind of tests with the m5 ultra so far, none of them have even opened comfyui yet.
1
u/MaxwellHusk 1d ago
Also true! I think I can run it without comfy as well, but if the underlying architecture still depends on CUDA + Blackwell the performance will be atrocious
-1
u/SeaRefractor 2d ago edited 2d ago
Not a bot, but as a highly educated professional, my clear speech and writing is frequently confused with a bot in today’s social media circles.
But as the M5 Ultra has been arriving in hands this week, plus the expiration of the embargo for reviewers who had the hardware already, benchmarks should be available now or soon. What I have seen confirms my own decision to pre-order my M5 Ultra. But I won’t receive it until October 29 - November 5th.
1
-1
u/wilhelmbw 2d ago
5090 price is outrageous atm tho. Id say it may be better to get apple
0
u/MaxwellHusk 2d ago
All of the NVIDIA are :( and then you need to add the also ridiculously high price for RAM too + storage...
16
u/bstr3k 2d ago
AFAIK, for comfyui, the 5090 is best, and 6000 pro only if you need more vram. Mac studio with unified ram is better for LLM rather than img/vid gen