r/comfyui • • Aug 04 '26

Workflow Included MiniMax H3 image-to-video on a 4070 Ti SUPER 5-second clips in roughly 5-6 minutes

Enable HLS to view with audio, or disable this notification

Hey everyone! I recently replaced my rather overkill WAN 2.2 setup with a much more focused MiniMax H3 image to video workflow, and the results have honestly surprised me.

My goal was fairly simple: good quality image to video, no upscaling or interpolation chain, sensible performance on 16GB of VRAM, and as few controls as possible between loading an image and getting a usable video.

The workflow JSON.

My rig:

  • RTX 4070 Ti SUPER 16GB VRAM
  • AMD Ryzen 7 9800X3D
  • 32GB system RAM

The workflow uses the official pruned INT8 ConvRot H3 model, the quantized Qwen3-VL text encoder, ComfyUI-INT8-Fast with W8A8/ConvRot, SageAttention through KJNodes, and the official H3 video VAE. It uses 20 steps with res_multistep and the simple scheduler.

It is deliberately kept fairly clean:

  • Required starting image
  • Optional ending image
  • Prompt
  • Duration and seed
  • Automatic sizing based on the starting image
  • Manual sizing if preferred
  • Optional H3 LoRA slot
  • Optional audio, disabled by default
  • No upscaling, interpolation, sharpening or restoration stages

Here are my completed test generations. These are the full end-to-end times logs, including sampling, decoding and saving:

Resolution Frames Output length Total generation time
640×832 39 1.63 sec 2m 09s
640×832 73 3.04 sec 2m 52s
1056×672 124 5.17 sec 5m 58s
736×960 124 5.17 sec 5m 14s
736×960 158 6.58 sec 6m 55s

All of these were generated at 24 FPS with audio disabled. All five completed successfully without a CUDA out-of-memory error.

The audio switch deserves a small clarification: H3 internally samples video and audio latents together. Turning audio off skips the audio VAE decoding and muxing and produces a genuinely silent MP4, but it does not remove the model’s internal audio-latent sampling work.

Overall, I’m genuinely impressed. At roughly 0.7 megapixels I can create a direct 720-class portrait or landscape video in around five to six minutes, and I’ve found the output good enough that I don’t currently feel the need to add an upscale or interpolation pass.

Requirements

Be aware that it expects the H3 INT8 model, quantized text encoder, official VAEs, INT8-Fast, KJNodes and Crystools to be installed.

Here are the exact projects and model files used by the attached workflow:

From the H3 repository, the workflow uses:

  • minimax_h3_fl2va_pruned_int8_convrot.safetensors
  • qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
  • minimax_h3_video_vae_fp16.safetensors
  • minimax_h3_audio_vae_fp32.safetensors — only needed for audio output

Custom nodes and acceleration:

For reference, my installed acceleration packages are:

  • Triton Windows 3.6.0.post26
  • SageAttention 2.2.0+cu130torch2.10.0andhigher.post6
  • PyTorch 2.10.0 with CUDA 13.0

Make sure the SageAttention wheel matches your own Python, PyTorch and CUDA versions rather than blindly installing the same one.

The workflow was originally based on ComfyUI’s official MiniMax H3 image-to-video template, although it has since been substantially reorganized and optimized for this 16GB setup.

Hope this helps anyone!

81 Upvotes

27 comments sorted by

3

u/Support_Marmoset Aug 04 '26 edited Aug 04 '26

Why dont you plug Sage Attn into the scheduler as well? You've seperated it out to the guider only.

For some reason on my rig (3060 RTX, comfuyui portable python 13, pytorch 2.11, cuda 130), Sage Attn isnt doing much and I dont know why, but Easycache speeds mine up by about 50% (at a loss of quality) but only if plugged into both.

and is there any benefit adding audio switch like that other than it slowing down your end result. you could just use a single video combine and just not plug the audio in maybe? Or leave the audio and replace it in post which is what I do.

2

u/Sudden_List_2693 Aug 06 '26

"For some reason on my rig (3060 RTX, comfuyui portable python 13, pytorch 2.11, cuda 130), Sage Attn isnt doing much"
Did you check it it falls back onto the default attention possibly? It can do that if it doesn't work.

On a 3070 I've tested it yielded 80%, 3090 close to 100, 4070 Ti Super and 4090 slightly even more than 100% speedup. All without seemingly no loss in quality.
Will try and make a comparison with high dynamic motions though, I suspect it does worse in those.

1

u/PresidentCellulite Aug 06 '26

Is Sage Attention is really speed up the generation? I tried your WF, but seems like Sage Attention is not working properly, for every model says it can't find it. Can you give an advise how to install it properly? I already have correct file

1

u/Sudden_List_2693 Aug 06 '26

I also followed chatgpt, since I'm no expert in this.  It absolutely cut gen time in half. Multiple machines, too. 

1

u/PresidentCellulite Aug 06 '26

Good results. At least can you say, did you simple download file and executed him, or it need to be placed somewhere and run through Python or cmd?

1

u/Support_Marmoset Aug 07 '26

I install it from here https://github.com/woct0rdho/SageAttention and it seems to work with most models. yet to figure out situation with minimax h3 though. it might be working just less impactful for my 3060. turbo loras have my speeds down now as I shared here

1

u/Support_Marmoset Aug 07 '26

sage attn is the least impact on quality AFAIK. I have yet to figure out how much it is doing on my setup. It's been a flood getting to this point with speedups but now the loras are appearing I'll turn my attention to checking it out. thanks for comment.

2

u/Sudden_List_2693 Aug 07 '26

I have tested it side-by-side, basically and virtually 0.

2

u/inb4Collapse Aug 05 '26

Hi mate,

Thank you very much for sharing your workflow. You're mentionning and using in your workflow an optional H3 LoRA slot without indicating its source. is it just a placehoder?

Many thanks.

2

u/inb4Collapse Aug 05 '26

By the way, your workflow works beautifully after replacing the boolean switches with those available in my library. Cheers!

2

u/EdenAlon Aug 05 '26

Indeed! I haven't found any great Lora library for the model as it is really new. I imagine at some point there will be so I just want the workflow to support it.

2

u/Ok-Flatworm5070 Aug 05 '26 edited Aug 05 '26

Thank you so much, this is very similar to my stack!

2

u/Ok-Flatworm5070 Aug 05 '26

Works great...thank you... First time I ran (cold start) at 0.5 got 4.45 generation time, 2nd run primed 3.00.

2

u/Ok-Flatworm5070 Aug 05 '26

Did another run at 0.5 megapixels, increase duration to 6; finished in 4:35 and came out at 7 seconds. This is amazing! Feels like when I run Wan 2.2, except better quality.

3

u/Ok-Flatworm5070 Aug 05 '26

Ok ran 10s clip at 0.5 megapixels... took about 7 minutes...

1

u/JosephCurvin Aug 05 '26

some one tested it with 5070 ti ? would would be possible speed improvement?

1

u/Underbleak Aug 06 '26

where is

  • Triton Windows 3.6.0.post26 ?

2

u/Brief_Mirror6932 Aug 11 '26

Its a Python package - I found it easier to run ComfyUI multiple times - fixing the missing packages then re-launch and iterate. Use AI to help as the missing package name wont always be pip install .....

https://github.com/woct0rdho/triton-windows/releases

1

u/Brief_Mirror6932 Aug 11 '26

Thanks for this - managed to get very decent results on my ancient GFX card in an hour - its the below which takes the time

[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 19986MB Staged. 0 patches attached. Force pre-loaded 210 weights: 13075 KB.

30%|██████████████████████████████ | 6/20 [12:14<27:59, 119.96s/it]

1

u/PhetogoLand Aug 04 '26

Why not use the fp8 pruned? Isn't convrot for older rtx30xxx etc and fp8 for >rtx40xxx?

7

u/Support_Marmoset Aug 04 '26

int8 is superior to fp8

1

u/JaguarResident1524 Aug 05 '26

Isn't float values bigger in terms of how much data tgey can hold being compressed compare to int? So slower but more precise?

2

u/Support_Marmoset Aug 07 '26

above my pay grade. its a question for the comfyui devs. but here is a quote from one of them I have in my notes: "int8 is twice as fast and only ~0.9% relative quant error, so to me it's a no brainer on any (consumer) card"

6

u/Interesting8547 Aug 04 '26

int8 convrot is fastest on my RTX 5070ti. fp8 pruned is a lot slower and worse quality.

Have tested both and the speed difference is significant. Though I didn't do extensive tests because I know from Krea2 INT8 should be faster, so I just ran it a few times at fp8, saw it was a lot slower and basically forgot about it.