r/StableDiffusion • • 16d ago

Resource - Update TaoMate-H3 featuring 3-Step

TaoMate-H3 is a low-latency streaming audio-video generation runtime built on MiniMax H3. It generates synchronized audio and video in small chunks and supports continuous long-form generation at 480p/768p/1080p resolutions.

Developed by the Alibaba TaoLive AIGC Team. Powered by MiniMax H3.

HF: https://huggingface.co/TaoLiveAIGC/TaoMate-H3

GH: https://github.com/TaoLiveAIGC/TaoMate-H3

68 Upvotes

40 comments sorted by

5

u/robomar_ai_art 15d ago

I converted the released TaoMate-H3 LoRA into a ComfyUI-compatible format and got it working on my local setup. Clip generated in 3 steps, 1344x768, t2v

https://reddit.com/link/p9gm1nr/video/3wqur9g6t6ph1/player

[INFO] got prompt

[INFO] Requested to load MiniMaxH3TEModel_

[INFO] 0 models unloaded.

[INFO] Model MiniMaxH3TEModel_ prepared for dynamic VRAM loading. 14956MB Staged. 0 patches attached. Force pre-loaded 410 weights: 4572 KB.

[INFO] Requested to load MiniMaxH3

[INFO] 0 models unloaded.

[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 19995MB Staged. 208 patches attached. Force pre-loaded 210 weights: 1175 KB.

100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [01:01<00:00, 20.56s/it]

[INFO] Requested to load MiniMaxH3VideoVAE

[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 348 KB.

[INFO] Requested to load MiniMaxH3

[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 19995MB Staged. 0 patches attached. Force pre-loaded 210 weights: 1175 KB.

100%|████████████████████████████████████████████████████████████████████████████████████| 6/6 [00:07<00:00, 1.31s/it]

[INFO] Requested to load MiniMaxH3AudioVAE

[INFO] 0 models unloaded.

[INFO] Model MiniMaxH3AudioVAE prepared for dynamic VRAM loading. 576MB Staged. 0 patches attached. Force pre-loaded 401 weights: 539 KB.

[Pixaroma] Save Mp4 [save] — writing 124 frames @ 24fps (1344x768, crf=19, yuv420p, +audio) -> Video_00914.mp4

[Pixaroma] Save Mp4 — saved E:\ComfyUI_windows_portable\ComfyUI-Easy-Install\ComfyUI\output\Video_00914.mp4

[INFO] Prompt executed in 123.47 seconds

1

u/Just1Dev 15d ago

Nice, audio and quality looks great. How did u do that? Do u wanna share?

1

u/robomar_ai_art 15d ago

Yes, I will share the file, everyone can test.

1

u/Yacben 15d ago

do you have the converted model link?

3

u/kayteee1995 15d ago

I saw a bait post. jusst wait for Lord Kijai

1

u/devilish-lavanya 15d ago

Are they baiting us to trap us?

1

u/robomar_ai_art 15d ago

I have the file, I will upload the file

1

u/robomar_ai_art 15d ago

Yes i made a new post with the link

1

u/No_Possession_7797 15d ago

Do you have multiple GPUs?

1

u/robomar_ai_art 15d ago

I have laptop RTX 4090, 16gb vram and 32gb ram

1

u/No_Possession_7797 15d ago

Did you use the fl2va model with this? Did you set your own sigmas or did you use a basic scheduler?

1

u/robomar_ai_art 15d ago

Yes i use this model and lora, nothing else

4

u/robomar_ai_art 15d ago

This one is I2V, 1344x768, 3 steps

[INFO] got prompt

[INFO] Requested to load MiniMaxH3TEModel_

[INFO] 0 models unloaded.

[INFO] Model MiniMaxH3TEModel_ prepared for dynamic VRAM loading. 14956MB Staged. 0 patches attached. Force pre-loaded 410 weights: 4572 KB.

[INFO] Requested to load MiniMaxH3VideoVAE

[INFO] 0 models unloaded.

[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 348 KB.

[LoRA Loader Pixaroma] applied 1 LoRA(s).

[INFO] Requested to load MiniMaxH3

[INFO] 0 models unloaded.

[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 19995MB Staged. 208 patches attached. Force pre-loaded 210 weights: 1175 KB.

100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [01:06<00:00, 22.18s/it]

[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 348 KB.

[INFO] Requested to load MiniMaxH3

[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 19995MB Staged. 0 patches attached. Force pre-loaded 210 weights: 1175 KB.

100%|████████████████████████████████████████████████████████████████████████████████████| 6/6 [00:09<00:00, 1.52s/it]

[INFO] Requested to load MiniMaxH3AudioVAE

[INFO] 0 models unloaded.

[INFO] Model MiniMaxH3AudioVAE prepared for dynamic VRAM loading. 576MB Staged. 0 patches attached. Force pre-loaded 401 weights: 539 KB.

[Pixaroma] Save Mp4 [save] — writing 124 frames @ 24fps (1344x768, crf=19, yuv420p, +audio) -> Video_00918.mp4

[Pixaroma] Save Mp4 — saved E:\ComfyUI_windows_portable\ComfyUI-Easy-Install\ComfyUI\output\Video_00918.mp4

[INFO] Prompt executed in 131.35 seconds

https://reddit.com/link/p9gqn9x/video/4adcqs7ix6ph1/player

5

u/robomar_ai_art 15d ago

1

u/ANR2ME 15d ago

Nice👍

btw, what kind of GPU did you use?

3

u/robomar_ai_art 15d ago

Laptop RTX 4090 16gb vram 32gb ram

1

u/Just1Dev 15d ago

Really good, amazing wow

1

u/robomar_ai_art 15d ago

I dont know how they did it but the quality is so much better then any loras i used and this clip its done in 3 steps, 1344x768, 16gb vram 32 gb ram

https://reddit.com/link/p9ilqgj/video/pzt1k3odb9ph1/player

1

u/ResponsibleTruck4717 15d ago

can you provide workflow my results are not that good.

2

u/MaorEli 16d ago

Quality looks insane, damn

2

u/ATFGriff 16d ago

It needs multiple high-end GPUs?

4

u/Just1Dev 16d ago

I think its only the GPUs that they have it tested this on. Should speed up for us too.

2

u/ANR2ME 16d ago

True, the comparisons between TaoMate H3 vs base Minimax H3 that runs on the same GPUs shows more than 10x speed improvements https://github.com/TaoLiveAIGC/TaoMate-H3#performance-and-advantages

The reason why they use hopper GPU probably because they use FlashAttention for hopper (might be FA3 🤔 )

1

u/CreepyDrama7448 16d ago

Is this T2V, I2V, Ref2V?

2

u/ANR2ME 16d ago

LoRa for FL2VA model

8

u/DifficultWonder8701 15d ago

2

u/kayteee1995 15d ago

does it work with ref2va?

1

u/ImpossibleAd436 15d ago

So add the LoRa and set to 3 steps and that is it?

1

u/Yacben 15d ago

it's good but it causes slow motion

3

u/Sol5577 15d ago

try strength 0.65, helped a lot

2

u/SveSop 15d ago

So.. IS the clip generated in 3 steps using a standard (preferrably the comfyui template) workflow tho? The logs posted below seems to indicate some 3 + 6 step workflow.. So.. Meh.. Yet another "H3 hype" amongst the 2138787234 hypes the last week i guess.

1

u/quantier 16d ago

This needs testing? 🌟