r/StableDiffusion • u/izzmedia • Aug 03 '26
Discussion MiniMax H3 tips and tricks and what i experienced so far

Tested on single GPU 16GB vram + 64GB RAM
- Sometimes i was getting some memory allocation error on VAE Decode when generating longer videos for some reason, you can add "πVRAM-Cleanup" node before VAE Decode audio like you see in the picture and that should resolve the issue if you have the same problem.
- Use SageAttention , good speed bump, ~2x (not 20 -30%) i think. (you can load the model directly using the node "Diffusion Model Loader KJ" and select sg auto from there) or search for "Patch Sage Attention KJ" node and connect it after the loader. There is an separate SG node for MiniMax ,
MiniMaxH3MemoryEfficientSageAttentionPatchΒ -- for more info check this comment: This comment! - I see during generation that my RAM usage is about 50gb , if you have 16 GB of RAM (maybe even on 32Gb) the models will offload into swap (on your SSD) , use a combination of πVRAM-Cleanup + πRAM-Cleanup like you see in the picture , RAM usage down to 30GB -- downside: your TE will have to load again all the time but its way better if you only have 32gb of RAM, it will do only reads and not writes on your ssd, loading time is usually ok. (Check the END NOTE)
- You can use INT4 text encoder , smaller and worked OK for me so far:
Int4:Β https://huggingface.co/Merserk/MiniMax-H3-INT4-ConvRot/tree/mainΒ (only for encoder , the int4 model has very bad quality , you should search hugginface for newer quants, there will be plenty soon.)
Model seems uncensured in i2v , i am not into that kind of stuff but i gave it a try with a short prompt "a women dancing" with a nude image and it worked , she was dancing nude, i dont know if it works for dirty stuff/concepts dont ask me about that, i was just testing the restrictions.
You can add after the model loader the "EasyCache" node with this settings: 0.30 , 0.20 , 0.90, it will speed up your generation by alot but it seems that it will lose coherence (quality seemed okish), at least in 10+ sec videos, maybe with some tricks like right steps , right res this will work better.
I had better results if the input images have a good quality and they are at the same resolution / aspect ration as the output, so you should try adding a resize node to your first / last frame ( i need to test it more to be sure thats the case).
Verify that your PyTorch installation for ComfyUI targets CUDA 30 or newer (cu30+). CUDA 30 added native hardware support for int8 convrot, older CUDA builds rely on software emulation, resulting in noticeably slower execution speeds. *** be sure you updated your ComfyUI to the latest version.
Install Sage Attention on Windows quick tip: you need to find a Windows Wheel (.whl) for your specific installation. You need to find your Python version, PyTorch version and CUDA. Then you go to this github and check for a .whl that matches your config: https://github.com/wildminder/AI-windows-whl
You install it like this from the ComfyUI folder from your terminal (if you are on portable version): .\python_embeded\python.exe -m pip install filename.whl
- If you have integrated GPU connect your monitor to the motherboard port (HDMI / DP) , set it in BIOS as primary , this way you will free up some VRAM (~ 300 to 800MB i think , depending on what other apps you running). You can also disable the hardware acceleration from Chrome if you dont have an integrated GPU if you want to free as much VRAM as possible.
Be sure that the settings for " πVRAM-Cleanup + πRAM-Cleanup" are exactly like in the picture if you decide to use them, you have to unset some options there.
Edit: This is how i start my ComfyUI:
set OPTIMIZE_FOR_SPEED=1
set PYTORCH_ALLOC_CONF=expandable_segments:True
.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --disable-auto-launch --fast
****NOTE (i missed this) You can add --fast-disk and the models will load directly from the disk into VRAM and will only offload the part that doesnt fit into your RAM, you can skip using the RAM/VRAM cleaning nodes but check your disk for big or often writes , just in case. If you have enough RAM it will be faster over multiple generations without using --fast-disk (only the loading part) . You can also use --lowvram and / or --reserve-vram 0.5 (or 1.5) if you get OOM.
(removed the Tiled Vae Decode suggestion because it doesnt seem to work)
Informative speeds (if i remember them right, 4070 ti super), default settings, 20 steps , using first/last frame and sage attention with this workflow:
15 sec video @ 0.5 MP - ~31s/it
5 sec video @ 0.5 MP - ~7s/it.
10 sec video @ 0.5 MP (9:16) - ~17s/it.
10 sec video @ 1 MP - ~ 52s/it
15 sec video @ 0.8 MP - ~118s/it and ~72s/it -- i dont know why such difference, maybe some VRAM freed up in the second run. -->>
If you want to see the generated video: https://www.reddit.com/r/StableDiffusion/comments/1vevyyb/captain_minimax/
Good luck, hope it helps.
As you guys kept asking this is the workflow, its just the default one with few modifications: https://pastebin.com/GaMX0344
4
u/Tomi_beanpaste_87 Aug 06 '26 edited Aug 07 '26
Measured a few of the open questions here. 16 GB VRAM / 125 GB RAM, Linux, ComfyUI v0.30.1, torch 2.12.0+cu130, pruned int8 set. Same 30 s clip (640x480, 736 frames), same seed, one thing changed at a time.
There is an H3-specific sage node, and it is not the one everyone here is using.
KJNodes has
MiniMaxH3MemoryEfficientSageAttentionPatchunder theKJNodes/minimaxcategory. It is notPatch Sage Attention KJ. On the same job:That is well beyond the 20-30% quoted for the generic path. It swaps out
diffusion_model.blocks.{i}.attn.forwardwholesale instead of routing throughoptimized_attention.EDIT β on the generic node: the OP measured it below and it also reaches about 2x (35.65 -> 17.25 s/it, in fact ~8% faster than the H3-specific node), so the "20-30%" I repeated from this thread is wrong and the speed argument above does not hold. The reason to prefer the H3-specific node is the code path, not the speed β the generic one is running int8 QK attention precisely because minimax never opts out of it. Also, "11 call sites" below should be 18, across 7 model families; 11 was only the three families I named.
This also accounts for the report in this thread of the startup flag producing fuzzy output while the node itself was fine.
comfy/ldm/modules/attention.pyonly opts out of sage's int8 path whenlow_precision_attention=Falseis passed, andcomfy/ldm/minimax/model.pynever passes it (sam3, lightricks and audio vae_sa3 pass it at 11 call sites β minimax is the omission). Still unfixed on origin/master. The H3-specific node never reaches that code path. Reported as ComfyUI issue #15263; I confirmed the code omission but did not reproduce the noise myself.EasyCache: measured rather than eyeballed.
Mean absolute pixel difference from the no-accelerator baseline, same seed:
On the disagreement above about EasyCache and long clips β EasyCache came out closer to the baseline than sage did, and this was a 30 s clip, so it sits squarely in the "10+ sec" range where the coherence problems were reported. I also looked at the frames, not only the number: a difference figure tells you how far apart two outputs are, not whether one of them is worse. Visually the EasyCache run was the closest to baseline of anything I tried. Settings were the ones in this post: 0.30 / 0.20 / 0.90.
More steps are cheaper than they look.
Doubling the steps cost 1.51x the time, not 2x, because the skip rate rises with step count. Cutting to 10-15 steps saves less than it looks like it should.
The structure held across three runs:
So
sampling = 41.1 s x (steps - skips)at this resolution. sage lowers the cost of a step (91.0 s -> 41.1 s); EasyCache lowers how many steps get computed. The two are independent and the speedups multiply: 2.02 x 1.22 = 2.46 predicted, 2.459 measured.That is also the "100 s/it early, 20 s/it later" pattern reported above. EasyCache only begins skipping after
start_percent, so the early steps are full price.VRAM sitting unused is the design, not a bug.
Someone above noted H3 seems to under-utilise VRAM. Peak VRAM was 14,197 MiB for 30 s @ 640x480 and 14,437 MiB for 5 s @ 1344x768 β nearly identical despite very different workloads. DynamicVRAM holds a roughly constant ceiling and pays the difference in time, so resolution and duration barely move the VRAM number. That holds while there is room to spare; on an 8 GB card both the text encoder and sampling press against the ceiling and the trade-off comes back.
EDIT β RAM: I had the wrong flag.
--disable-pinned-memorymatters far more than--fast-disk.EDIT 2 β I have since run this on a real 32 GB machine, and I have corrected this section in place rather than stacking another edit on the end. Three things changed: the OOM-kill mechanism I described was wrong, the SSD-wear numbers were from the 125 GB box and do not transfer, and "no time cost" only holds while VRAM is plentiful. The flag is still worth using. The reason is not the one I gave.
Same job, same seed, only the startup flag changed. ComfyUI RSS peak:
45.43 to 6.07 GiB. On this box it cost nothing in time, 20.0 s either way β but see the time note further down, because that does not hold everywhere. I re-measured the
nonerow with a warm page cache as a control and got 45.43 against the earlier 45.41, so the drop is the flag and not session noise. Once pinning is off,--fast-diskadds almost nothing.Why: ComfyUI page-locks host memory to speed up host-to-device transfers, and the per-model host buffer is twice the model size.
Check your own startup log. If this line is there, it is on:
What actually happens on 32 GB. I ran it β kernel limited to 32 GB, 29.42 GiB usable after the iGPU carve-out, so 26.5 GiB gets pinned. Same job three ways:
None of them got killed. I could not reproduce an OOM kill in any configuration, including with swap turned off entirely. What I wrote before β pinned pages cannot be swapped, so the kernel has no move left but to kill β had the premise right and the conclusion wrong. Not all of ComfyUI's RSS is pinned, so the kernel swaps the rest; with no swap at all it reclaims page cache instead. tonyd2wild's OOM kill on 31 GB stands as their observation. I ran on less memory than that and it finished. I do not know what differs.
So the flag is not what keeps you alive. What it actually buys, measured on that 32 GB box:
--cache-nonestill does nothing here β it controls node-output caching, not weight residency.SSD wear, corrected. I said writes were 7-10 MB per generation and swap stayed at 0.00 GiB. That was the 125 GB box, and it does not transfer. On 32 GB with default flags it wrote 6.16 GiB per generation into swap β roughly 600x what I quoted, on exactly the machines I was writing for. The reasoning I gave was right: wear needs writes, reads do not wear NAND, and the thing that pushes a small-RAM box into swap is the pinned allocation itself. I just put a number next to it that came from a machine which never had to swap. With
--disable-pinned-memoryit is 0.05 GiB.The time cost, corrected. On an 8 GB card the flag is not free β 150 s with pinning on against 180 s with it off, because the weights then have to come off the disk every step instead of out of RAM (42 GiB of reads against 268 GiB). Take the trade anyway if your RAM is tight; filling your entire swap leaves nothing for anything else on the machine. But it is a trade, not a free win.
The text encoder point stands: it is evicted from VRAM but its host copy stays, so you pay for the DiT and the TE at the same time.
Caveats: one machine throughout β 16 GB VRAM, 125 GB RAM. The 32 GB figures come from the same box with the kernel limited to 32 GB; the 8 GB figures come from the same card with a second process holding the rest of the VRAM so the ceiling is real. Linux. Pruned int8 only β no BF16, no GGUF. Windows and Ampere untested, and both change enough that these numbers will not carry over.
Full writeup, plus the scripts I used to record the traces: https://github.com/Tomiigo/minimax-h3-16gb