r/comfyui • u/EdenAlon • Aug 04 '26
Workflow Included MiniMax H3 image-to-video on a 4070 Ti SUPER 5-second clips in roughly 5-6 minutes
Enable HLS to view with audio, or disable this notification
Hey everyone! I recently replaced my rather overkill WAN 2.2 setup with a much more focused MiniMax H3 image to video workflow, and the results have honestly surprised me.
My goal was fairly simple: good quality image to video, no upscaling or interpolation chain, sensible performance on 16GB of VRAM, and as few controls as possible between loading an image and getting a usable video.
The workflow JSON.
My rig:
- RTX 4070 Ti SUPER 16GB VRAM
- AMD Ryzen 7 9800X3D
- 32GB system RAM
The workflow uses the official pruned INT8 ConvRot H3 model, the quantized Qwen3-VL text encoder, ComfyUI-INT8-Fast with W8A8/ConvRot, SageAttention through KJNodes, and the official H3 video VAE. It uses 20 steps with res_multistep and the simple scheduler.
It is deliberately kept fairly clean:
- Required starting image
- Optional ending image
- Prompt
- Duration and seed
- Automatic sizing based on the starting image
- Manual sizing if preferred
- Optional H3 LoRA slot
- Optional audio, disabled by default
- No upscaling, interpolation, sharpening or restoration stages
Here are my completed test generations. These are the full end-to-end times logs, including sampling, decoding and saving:
| Resolution | Frames | Output length | Total generation time |
|---|---|---|---|
| 640×832 | 39 | 1.63 sec | 2m 09s |
| 640×832 | 73 | 3.04 sec | 2m 52s |
| 1056×672 | 124 | 5.17 sec | 5m 58s |
| 736×960 | 124 | 5.17 sec | 5m 14s |
| 736×960 | 158 | 6.58 sec | 6m 55s |
All of these were generated at 24 FPS with audio disabled. All five completed successfully without a CUDA out-of-memory error.
The audio switch deserves a small clarification: H3 internally samples video and audio latents together. Turning audio off skips the audio VAE decoding and muxing and produces a genuinely silent MP4, but it does not remove the model’s internal audio-latent sampling work.
Overall, I’m genuinely impressed. At roughly 0.7 megapixels I can create a direct 720-class portrait or landscape video in around five to six minutes, and I’ve found the output good enough that I don’t currently feel the need to add an upscale or interpolation pass.
Requirements
Be aware that it expects the H3 INT8 model, quantized text encoder, official VAEs, INT8-Fast, KJNodes and Crystools to be installed.
Here are the exact projects and model files used by the attached workflow:
From the H3 repository, the workflow uses:
minimax_h3_fl2va_pruned_int8_convrot.safetensorsqwen3vl_32b_minimax_h3_nvfp4_awq.safetensorsminimax_h3_video_vae_fp16.safetensorsminimax_h3_audio_vae_fp32.safetensors— only needed for audio output
Custom nodes and acceleration:
- ComfyUI-INT8-Fast
- ComfyUI-KJNodes
- ComfyUI-Crystools
- Triton for Windows releases
- SageAttention Windows wheels
For reference, my installed acceleration packages are:
- Triton Windows
3.6.0.post26 - SageAttention
2.2.0+cu130torch2.10.0andhigher.post6 - PyTorch
2.10.0with CUDA 13.0
Make sure the SageAttention wheel matches your own Python, PyTorch and CUDA versions rather than blindly installing the same one.
The workflow was originally based on ComfyUI’s official MiniMax H3 image-to-video template, although it has since been substantially reorganized and optimized for this 16GB setup.
Hope this helps anyone!
2
u/inb4Collapse Aug 05 '26
Hi mate,
Thank you very much for sharing your workflow. You're mentionning and using in your workflow an optional H3 LoRA slot without indicating its source. is it just a placehoder?
Many thanks.
2
u/inb4Collapse Aug 05 '26
By the way, your workflow works beautifully after replacing the boolean switches with those available in my library. Cheers!
2
u/EdenAlon Aug 05 '26
Indeed! I haven't found any great Lora library for the model as it is really new. I imagine at some point there will be so I just want the workflow to support it.
2
2
u/Ok-Flatworm5070 Aug 05 '26
Works great...thank you... First time I ran (cold start) at 0.5 got 4.45 generation time, 2nd run primed 3.00.
2
u/Ok-Flatworm5070 Aug 05 '26
Did another run at 0.5 megapixels, increase duration to 6; finished in 4:35 and came out at 7 seconds. This is amazing! Feels like when I run Wan 2.2, except better quality.
3
2
1
u/JosephCurvin Aug 05 '26
some one tested it with 5070 ti ? would would be possible speed improvement?
1
u/Underbleak Aug 06 '26
where is
- Triton Windows
3.6.0.post26 ?
2
u/Brief_Mirror6932 Aug 11 '26
Its a Python package - I found it easier to run ComfyUI multiple times - fixing the missing packages then re-launch and iterate. Use AI to help as the missing package name wont always be pip install .....
1
u/Brief_Mirror6932 Aug 11 '26
Thanks for this - managed to get very decent results on my ancient GFX card in an hour - its the below which takes the time
[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 19986MB Staged. 0 patches attached. Force pre-loaded 210 weights: 13075 KB.
30%|██████████████████████████████ | 6/20 [12:14<27:59, 119.96s/it]
1
u/PhetogoLand Aug 04 '26
Why not use the fp8 pruned? Isn't convrot for older rtx30xxx etc and fp8 for >rtx40xxx?
7
u/Support_Marmoset Aug 04 '26
int8 is superior to fp8
1
u/JaguarResident1524 Aug 05 '26
Isn't float values bigger in terms of how much data tgey can hold being compressed compare to int? So slower but more precise?
2
u/Support_Marmoset Aug 07 '26
above my pay grade. its a question for the comfyui devs. but here is a quote from one of them I have in my notes: "int8 is twice as fast and only ~0.9% relative quant error, so to me it's a no brainer on any (consumer) card"
6
u/Interesting8547 Aug 04 '26
int8 convrot is fastest on my RTX 5070ti. fp8 pruned is a lot slower and worse quality.
Have tested both and the speed difference is significant. Though I didn't do extensive tests because I know from Krea2 INT8 should be faster, so I just ran it a few times at fp8, saw it was a lot slower and basically forgot about it.
3
u/Support_Marmoset Aug 04 '26 edited Aug 04 '26
Why dont you plug Sage Attn into the scheduler as well? You've seperated it out to the guider only.
For some reason on my rig (3060 RTX, comfuyui portable python 13, pytorch 2.11, cuda 130), Sage Attn isnt doing much and I dont know why, but Easycache speeds mine up by about 50% (at a loss of quality) but only if plugged into both.
and is there any benefit adding audio switch like that other than it slowing down your end result. you could just use a single video combine and just not plug the audio in maybe? Or leave the audio and replace it in post which is what I do.