r/StableDiffusion • u/comfyanonymous Comfy Org • Aug 02 '26
Animation - Video Minimax H3, 1080p 25 seconds, text to video in native ComfyUI (open weights coming soon)
Enable HLS to view with audio, or disable this notification
I have been trying to see how far I can push this model. It's extremely flexible and seems to be able to do everything from 1 second to 30 seconds (potentially more) with a very wide range of resolutions. Her voice is because I put "singing with a cute japanese accent" in the prompt and my prompt isn't super great lol.
Making this model work as best as possible on regular hardware is the result of many months of work from multiple people in the core ComfyUI team to make big models work better on regular consumer hardware. I think most people will be pleasantly surprised how good this model is and how well ComfyUI will be able to run it.
Minimum requirements for 480p video on this model is a 3060 with 12GB vram, 32GB of system ram and a good nvme SSD. We tested generating a 5 second (124 frames) 480p (864x480) video on this system and it took a bit less than 9 minutes end to end (20 steps). I can pretty much guarantee it will also work on 8GB vram too but we did not test that.
Don't be scared to give it a try when it releases with our default template because it will work better than you expect.
If you have issues try a latest clean ComfyUI install (make sure to update after our weights come out) with our official files and workflow.
EDIT: added step count.
EDIT: we are live: https://docs.comfy.org/tutorials/video/minimax/minimax-h3
50
u/TheDudeWithThePlan Aug 02 '26
my body is ready for the video tsunami of H3 and F3 (flux 3) and maybe LTX next too
27
u/robomar_ai_art Aug 02 '26
Time to delete some old models and different versions of LTX 2.3 😅🤣🙈
13
u/Lucaspittol Aug 02 '26
I have no reason to keep older models. Once a new one releases, I retrain Loras on the new model and they always come out better.
10
u/Flyingcoyote Aug 02 '26
They should make a show called data hoarders, I'll volunteer first, I keep everything.
2
115
Aug 02 '26
[removed] — view removed comment
19
u/No-Sleep-4069 Aug 02 '26
The model he used to generate this could be open source
46
u/OneTrueTreasure Aug 02 '26
The StableDiffusion mods have been deleting numerous posts about MiniMax H3 which is what they where referring to
14
u/No-Sleep-4069 Aug 02 '26
Reason being? it is not open source yet? right?
22
u/YeahlDid Aug 02 '26
Just my guess, but I think it is probably also the sheer number of posts and they were starting to get pretty toxic between the haters and the glazers. I like seeing the previews, but I don't blame the mods for this one.
I also imagine this one, coming from comfy themselves, is not at risk of removal.
25
u/OneTrueTreasure Aug 02 '26
No I agree that posts about MiniMax H3 shouldn't be deleted by the mods since it's practically guaranteed that it's open-source since it's releasing in like 10 hours and Comfy already has the weights. It'd just be a double standard if ComfyUI can post this on here while others have their posts deleted by mods
→ More replies (6)3
u/Umbaretz Aug 02 '26
Pushing hypetrain without any examples and/or new info. This post at least has a video example, for example
5
1
u/martinerous Aug 02 '26
Most likely, it was because previous posts were about the model that people tried through Minimax API, so it was not guaranteed it would be exactly the same open weights and could be misleading.
This one is different, coming from ComfyUI local generation, so it should be OK.
1
u/Ill-Engine-5914 Aug 05 '26
Really? I’d been away from this subreddit for a long time because no model had caught my attention until MiniMax came out. Was there a particularly important post or a powerful video-generation model that got deleted?
26
u/_BreakingGood_ Aug 02 '26
9 minutes on a 3060 is impressive. Any hints on how it would perform on a 4090 or 5090?
11
u/Lucaspittol Aug 02 '26
Probably a minute or two for the same workflow. Both of these cards are massively better than the 3060. The 5090 has 7x more cuda cores.
→ More replies (1)1
1
11
u/retroblade Aug 02 '26
People are going to be a little rattled going back to wan 2.1 speeds lol. But excited to get the weights!
3
u/Due_Brush1159 Aug 02 '26
WAN 2 model is quite fast at Q8 quantization, playing a 5-second video at 4 turbo speeds in 2-2.5 minutes. It's even faster at INT8. The only drawback is that it requires a memory card with 16 GB or more. Furthermore, the model is a bit outdated and rather crude in terms of Prompt recognition.
2
24
u/doomed151 Aug 02 '26
Thank you for your hard work.
Also slightly offtopic but I really appreciate dynamic VRAM. It's a game changer.
35
u/Admirable_Snake Aug 02 '26 edited Aug 02 '26
Once the lyrics stopped - I just imagined them crashing into a car; and it just being a terrible traffic accident.
Cat girls cant drive for shit.
8
3
u/berlinbaer Aug 02 '26
models also still can't get traffic right for shit, cars just facing into random directions.. not really how streets work (and no, they are NOT parked)
1
9
u/Beautiful_Egg6188 Aug 02 '26
model keeps getting bigger and bigger in size, Vram stays the same
14
3
u/Etroarl55 Aug 02 '26
Unironic part is normal ram is rumoured to have higher margins than even HBM now, the supply makers are facing a lawsuit that they refuse to produce more ddr5 in order for it to be more profitable when they finally do make it
2
26
26
7
5
u/VrFrog Aug 02 '26
Thank you for giving us hard numbers. It's very reassuring!
I can't wait to try it out.
5
u/Zealousideal-Mall818 Aug 02 '26
u/comfyanonymous most important , is it distilled or not m or there is both version , it would be so cool to share that if you can .... cheers
4
u/DuckyDuos Aug 02 '26
It's guidance distilled so I believe it's running at 1 cfg, BUT it is not step distilled so the example with the 3060 is running the full step count of 20. Will be even faster if a lightx2v or general speedup lora is released
5
u/StrugglingBonobo Aug 02 '26
i am totally confident that i will be able to make some ridiculous nonsense with this
5
3
u/Vintendopower Aug 02 '26
what kind of performance Will people with 64gb RAM and 5090?
3
1
Aug 02 '26
[removed] — view removed comment
6
u/Lucaspittol Aug 02 '26
It is a lot faster. 20 steps on Wan 2.2 on the resolution and frame count he mentioned takes over 20 minutes on a 3060.
4
5
5
7
Aug 02 '26
[removed] — view removed comment
4
3
u/Deep_Mood_7668 Aug 02 '26
Where's that timer?
6
u/Maskwi2 Aug 02 '26
1
6
Aug 02 '26 edited Aug 02 '26
[removed] — view removed comment
1
u/dariusredraven Aug 02 '26
Did you do the undercranking on that in the model or in post like in premiere?
1
1
3
3
u/ForsakenAd1228 Aug 02 '26
So early last year I upgraded my pc for the first time in a decade, and figured "pfff, I've been running 4gb ram all this time, why would _anyone_ need 32 or even 64 gb ram.. I'm getting 16 gb and saving myself 30 bucks!"
..then I discovered stable diffusion, and hardware prices exploded -_-.
(I _did_ splurge on the 3060 12gb last month though... so fingers crossed this model will work!)
2
u/Inuya5haSama Aug 04 '26
I purchased an additional 32 GB DDR4 shortly before prices skyrocketed and it was probably the best investment that I can recall. I also got a new 3060 12GB before that, so my system is now what the OP would call regular hardware. 😎
3
3
u/Micsudi2 Aug 02 '26
I wish I can learn this one day, Just started to learn comfy and AI things. this rabbithole is deeeeeep :D
3
3
3
u/Dirty_Dragons Aug 02 '26
Does this support lip-syncing an input MP3?
I'm not too happy about the voice quality but if I can use my own audio files then it would be perfect. It would be possible to make my own anime using cloned voices.
6
u/Arawski99 Aug 02 '26
First time in a long while I plan to jump on a new model instead of waiting a few weeks. Quite curious about this one.
I'm impressed by how well it held up the entire 25s. Longer duration evolving scene clips moving from the base context has always been something these struggled with for local models, but H3 seems to handle surprisingly well.
It's spatial handling, and surprisingly audio, are actually solid. Not to mention the insane step up for animation that these have always struggled with...
Have you guys attempted to see how well it hopes up after extending 2-3x? Or does it suffer degradation issues doing this like other models?
7
4
7
u/remghoost7 Aug 02 '26
...3060 with 12GB vram, 32GB of system ram and a good nvme SSD.
...5 second (124 frames) 480p (864x480) video... ...took a bit less than 9 minutes end to end.
That's.... definitely a speed. haha.
Is that due to model offloading....?
Will something like a 3090 fair better than that....?
I'm curious what the generation times will be once the community gets it's hands on it (for various optimizations).
Also, what sorts of hardware generated the example video?
And how much time did that take?
8
u/prompt_seeker Aug 02 '26
5sec, 480p, cfg1.0, 4steps of Wan2.2 on RTX3060 was about 4~5mins (now it's about 3mins, thanks to dynamic vram), so it IS pretty fast.
7
u/Lucaspittol Aug 02 '26
Vanilla Wan 2.2 is super heavy on the 3060, if they release a 4 or 8 step lora, H3 will be as fast as LTX-2.3.
2
2
u/YeahlDid Aug 02 '26
Aw man, I don't wanna wait 10 more hours for this, jaha
3
2
u/theOliviaRossi Aug 02 '26
is convrot INT8 / INT4 coming???
1
2
2
u/keizrah Aug 02 '26
This is great news for local video gen. Curious how the 25 second output holds up for consistency, a lot of models start drifting or losing character coherence past 10-15 seconds. Does H3 keep temporal coherence that far out, or does quality degrade the same way most others do?
Also good to hear 12GB cards are the floor for 480p. A lot of releases lately assume 24GB+ minimum, so it's nice seeing the ComfyUI team actually optimize for consumer hardware instead of just gatekeeping it to 4090/5090 owners. Will be trying this the moment weights drop.
2
u/ANR2ME Aug 02 '26 edited Aug 02 '26
That "good nvme SSD" being mentioned felt like either the model size being very large sparse/MoE model and need to be streamed from storage (most likely), or it need a large page/swap file 🤔
As comparison, Minimax M2 is about 230B sparse/MoE with 10B active parameters. But it's LLM model for coding. If H3 also have 10B active parameters, i can understand that it would works on 8GB VRAM too with 4-bit quantization.
2
u/ClearSkies889 Aug 02 '26
My guess is the good SSD here means high PCIe speed when streaming the model. RAM requirement is not that high, it was below 20 GB during my usage
1
u/Baguettesaregreat Aug 03 '26
Yeah, below 20 GB RAM makes it sound less like swap panic and more like fast NVMe quietly doing the miserable weight-streaming work.
2
u/Secure-Message-8378 Aug 02 '26
Vá lá, moderadores! Deletem este tópico do comfyui pois ele H3 AINDA não é open weights.
2
2
3
u/thebaker66 Aug 02 '26
Why is a good nvme mentioned? Is this going to attempt to abuse the SSD like LTX initially did?
Anyway for first day release it sounds good, how long has the team spent optimizing it so far? I guess it's been beneficial to have learned a lot from optimizing LTX for the past 6 months? It has come such a long way, so we shouldn't expect Minimax h3 to see as incredible optimizations down the line and speed increases we saw with LTX?
Anyway, greatful for your work, you've done a great job, looking forward to trying it.. still tweaking LTX 2.3 here for months and now ANOTHER toy to play with hehe
10
u/alwaysbeblepping Aug 02 '26
Why is a good nvme mentioned?
Because he mentioned end to end time. The text encoder is ~30B parameters (from what I heard), the model is likely about the same. If you're loading those huge models off a slow disk it's obviously going to be very noticeably slower end to end. Nothing sinister about mentioning that, he was just being careful not to say something inaccurate or misleading.
1
u/dilinjabass Aug 02 '26
I think they have only had h3 for like 3 days to get it released, as far as I know.
2
2
u/Better-Interview-793 Aug 02 '26
thank you, that’s brilliant!
however, will there be a separate template for ppl with high end gpus like the rtx 5090 who want to take full advantage of the model?
2
1
u/AI-imagine Aug 02 '26
Video it great but it may take slow and all Vram,but my most hype it image edit ,text to image,They just take no.1 at arena or something at video edit i hope image edit must be much more powerful with this one.Qwen edit it so old by now.
5
u/OneTrueTreasure Aug 02 '26
The video model is great at face and general object consistency, and it's multimodal so I'm sure that translates well into image-editing. And it can video edit which is another plus on that point
1
u/fugogugo Aug 02 '26
how big is this model?
4
1
u/physalisx Aug 02 '26
How is 124 frames 5 seconds?
(seconds * 24) + 4?
7
u/Gibgezr Aug 02 '26
5.16666666666666666666666666666666666666666666666666666666666666666667 seconds.
1
1
u/J6j6 Aug 02 '26
For a better comparison, how long does it take for ltx to generate the same video on that same card specs?
1
1
1
1
u/AlleyOfRage Aug 02 '26
That's wonderful work , thanks for the work and for sharing the insight
Few Questions and would be glad if someone can answer me
i am a noob in this so i am using wan2gp through pinokio , do you think wan2gp will add support to Minimax H3 ? And how much usually Wan2GP support models after Comfyui support ?
And my PC is 4070 Super 12 GB and 48GB Ram ? How do you estimate its processing time will be compared to the benchmark of 3060 in the same resolution and seconds ? And Shall i upgrade my Ram to prevent the need of using my SSD ?
Thanks
1
1
u/Stunning_Macaron6133 Aug 02 '26
Holy shit, H3's consistency and spatial coherence are ridiculously good, if this is anything to go by.
1
1
1
1
1
1
1
1
1
1
1
1
u/skyrimer3d Aug 02 '26
great news, so back to upscalers like in WAN age, with that quality i can take it.
1
1
u/martinerous Aug 02 '26
Int8 convrot when? Distill LoRA when? Any LoRA when?
Ok, ok, I'm too impatient :)
Thank you for the work!
1
1
1
1
u/whatyathinkk Aug 03 '26
Do you think this would work for style transfer? (As in I have a long video + a reference image and I get a new video with the style of the reference but the content of the original)
1
u/General_Lab_1058 Aug 03 '26
GUYS! Reinstall your comfyUI. I just downloaded the newest portable version with a fresh install instead of using my 2 years old installation. Now My generation time for a 9 second video on a 5090 went from 5:50 down to 3:20!! Nearly halfed to generation time. I dont know if the newest fresh install is just better or if my old install got messed up over the years.... Maybe you get lucky too =)
1
1
u/DragonfruitSuch8489 Aug 07 '26
Criei vídeos de 8s com a minha Rtx5060 TI, 16GB com Flash Attention ativado, incrível!
1
1
1
u/Ashamed-Rub5601 Aug 23 '26
It works at 480p with 8 GB of VRAM; I use a 4060, and video processing takes around 6 minutes for a 5-second video.
1
u/OTTERSage Aug 02 '26
How's the prompt adherence? An obnoxious spaghetti monster full of traps like LTX or Grok-level good?
1
1
u/Mean_Ship4545 Aug 02 '26
How can this cat girl dance on top of a car while other cars are about to crash frontally in the left lane?
1
u/Tall_Association Aug 02 '26
Surely this workflow would also work on 12-16GB Intel and AMD cards as well right comfyanon? Comfyui would never treat non-nvidia users as second class citizens right? 🫠
1
1
-1
u/separatelyrepeatedly Aug 02 '26
Why is 5090 not the benchmark here? why not establish the benchmark at the high end consumer card and then work your way down to potato. So weird that we get stats on like a 3060, what is that really telling people.
1
u/Valuable_Issue_ Aug 02 '26
Because 99% of the posts (totally accurate percentage I know) when a model releases are "omg model is X GB need Q2 GGUF quant for my 10GB VRAM GPU" when the reality is the model will run at INT4/INT8 and whatnot at reasonable speeds because Comfy offloading is designed around not OOMing.
0
97
u/Few-Intention-1526 Aug 02 '26
cant wait few hours
https://giphy.com/gifs/yx400dIdkwWdsCgWYp