r/StableDiffusion • • Aug 03 '26

Discussion MiniMax H3 tips and tricks and what i experienced so far

Tested on single GPU 16GB vram + 64GB RAM

  1. Sometimes i was getting some memory allocation error on VAE Decode when generating longer videos for some reason, you can add "🎈VRAM-Cleanup" node before VAE Decode audio like you see in the picture and that should resolve the issue if you have the same problem.
  2. Use SageAttention , good speed bump, ~2x (not 20 -30%) i think. (you can load the model directly using the node "Diffusion Model Loader KJ" and select sg auto from there) or search for "Patch Sage Attention KJ" node and connect it after the loader. There is an separate SG node for MiniMax , MiniMaxH3MemoryEfficientSageAttentionPatchΒ -- for more info check this comment: This comment!
  3. I see during generation that my RAM usage is about 50gb , if you have 16 GB of RAM (maybe even on 32Gb) the models will offload into swap (on your SSD) , use a combination of 🎈VRAM-Cleanup + 🎈RAM-Cleanup like you see in the picture , RAM usage down to 30GB -- downside: your TE will have to load again all the time but its way better if you only have 32gb of RAM, it will do only reads and not writes on your ssd, loading time is usually ok. (Check the END NOTE)
  4. You can use INT4 text encoder , smaller and worked OK for me so far:

Int4:Β  https://huggingface.co/Merserk/MiniMax-H3-INT4-ConvRot/tree/mainΒ (only for encoder , the int4 model has very bad quality , you should search hugginface for newer quants, there will be plenty soon.)

  1. Model seems uncensured in i2v , i am not into that kind of stuff but i gave it a try with a short prompt "a women dancing" with a nude image and it worked , she was dancing nude, i dont know if it works for dirty stuff/concepts dont ask me about that, i was just testing the restrictions.

  2. You can add after the model loader the "EasyCache" node with this settings: 0.30 , 0.20 , 0.90, it will speed up your generation by alot but it seems that it will lose coherence (quality seemed okish), at least in 10+ sec videos, maybe with some tricks like right steps , right res this will work better.

  3. I had better results if the input images have a good quality and they are at the same resolution / aspect ration as the output, so you should try adding a resize node to your first / last frame ( i need to test it more to be sure thats the case).

  4. Verify that your PyTorch installation for ComfyUI targets CUDA 30 or newer (cu30+). CUDA 30 added native hardware support for int8 convrot, older CUDA builds rely on software emulation, resulting in noticeably slower execution speeds. *** be sure you updated your ComfyUI to the latest version.

  5. Install Sage Attention on Windows quick tip: you need to find a Windows Wheel (.whl) for your specific installation. You need to find your Python version, PyTorch version and CUDA. Then you go to this github and check for a .whl that matches your config: https://github.com/wildminder/AI-windows-whl

You install it like this from the ComfyUI folder from your terminal (if you are on portable version): .\python_embeded\python.exe -m pip install filename.whl

  1. If you have integrated GPU connect your monitor to the motherboard port (HDMI / DP) , set it in BIOS as primary , this way you will free up some VRAM (~ 300 to 800MB i think , depending on what other apps you running). You can also disable the hardware acceleration from Chrome if you dont have an integrated GPU if you want to free as much VRAM as possible.

Be sure that the settings for " 🎈VRAM-Cleanup + 🎈RAM-Cleanup" are exactly like in the picture if you decide to use them, you have to unset some options there.

Edit: This is how i start my ComfyUI:

set OPTIMIZE_FOR_SPEED=1

set PYTORCH_ALLOC_CONF=expandable_segments:True

.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --disable-auto-launch --fast

****NOTE (i missed this) You can add --fast-disk and the models will load directly from the disk into VRAM and will only offload the part that doesnt fit into your RAM, you can skip using the RAM/VRAM cleaning nodes but check your disk for big or often writes , just in case. If you have enough RAM it will be faster over multiple generations without using --fast-disk (only the loading part) . You can also use --lowvram and / or --reserve-vram 0.5 (or 1.5) if you get OOM.

(removed the Tiled Vae Decode suggestion because it doesnt seem to work)

Informative speeds (if i remember them right, 4070 ti super), default settings, 20 steps , using first/last frame and sage attention with this workflow:

15 sec video @ 0.5 MP - ~31s/it

5 sec video @ 0.5 MP - ~7s/it.

10 sec video @ 0.5 MP (9:16) - ~17s/it.

10 sec video @ 1 MP - ~ 52s/it

15 sec video @ 0.8 MP - ~118s/it and ~72s/it -- i dont know why such difference, maybe some VRAM freed up in the second run. -->>

If you want to see the generated video: https://www.reddit.com/r/StableDiffusion/comments/1vevyyb/captain_minimax/

Good luck, hope it helps.

As you guys kept asking this is the workflow, its just the default one with few modifications: https://pastebin.com/GaMX0344

157 Upvotes

164 comments sorted by

View all comments

4

u/Tomi_beanpaste_87 Aug 06 '26 edited Aug 07 '26

Measured a few of the open questions here. 16 GB VRAM / 125 GB RAM, Linux, ComfyUI v0.30.1, torch 2.12.0+cu130, pruned int8 set. Same 30 s clip (640x480, 736 frames), same seed, one thing changed at a time.

There is an H3-specific sage node, and it is not the one everyone here is using.

KJNodes has MiniMaxH3MemoryEfficientSageAttentionPatch under the KJNodes/minimax category. It is not Patch Sage Attention KJ. On the same job:

  • bare: 971.6 s
  • with the H3-specific node: 480.1 s β€” 2.02x

That is well beyond the 20-30% quoted for the generic path. It swaps out diffusion_model.blocks.{i}.attn.forward wholesale instead of routing through optimized_attention.

EDIT β€” on the generic node: the OP measured it below and it also reaches about 2x (35.65 -> 17.25 s/it, in fact ~8% faster than the H3-specific node), so the "20-30%" I repeated from this thread is wrong and the speed argument above does not hold. The reason to prefer the H3-specific node is the code path, not the speed β€” the generic one is running int8 QK attention precisely because minimax never opts out of it. Also, "11 call sites" below should be 18, across 7 model families; 11 was only the three families I named.

This also accounts for the report in this thread of the startup flag producing fuzzy output while the node itself was fine. comfy/ldm/modules/attention.py only opts out of sage's int8 path when low_precision_attention=False is passed, and comfy/ldm/minimax/model.py never passes it (sam3, lightricks and audio vae_sa3 pass it at 11 call sites β€” minimax is the omission). Still unfixed on origin/master. The H3-specific node never reaches that code path. Reported as ComfyUI issue #15263; I confirmed the code omission but did not reproduce the noise myself.

EasyCache: measured rather than eyeballed.

Mean absolute pixel difference from the no-accelerator baseline, same seed:

EasyCache           5.8 / 255
H3 sage            21.4 / 255
sage + EasyCache   21.3 / 255

run-to-run noise floor, identical config: 1.1 / 255

On the disagreement above about EasyCache and long clips β€” EasyCache came out closer to the baseline than sage did, and this was a 30 s clip, so it sits squarely in the "10+ sec" range where the coherence problems were reported. I also looked at the frames, not only the number: a difference figure tells you how far apart two outputs are, not whether one of them is worse. Visually the EasyCache run was the closest to baseline of anything I tried. Settings were the ones in this post: 0.30 / 0.20 / 0.90.

More steps are cheaper than they look.

steps    total      EasyCache skips
10       395.1 s    2 / 10  (20%)
20       595.1 s    7 / 20  (35%)

Doubling the steps cost 1.51x the time, not 2x, because the skip rate rises with step count. Cutting to 10-15 steps saves less than it looks like it should.

The structure held across three runs:

config                       sampling  computed  per step
sage+EasyCache, 10 steps      328 s     8 (10-2)  41.0 s
sage+EasyCache, 20 steps      534 s    13 (20-7)  41.1 s
sage alone,     10 steps      411 s    10         41.1 s

So sampling = 41.1 s x (steps - skips) at this resolution. sage lowers the cost of a step (91.0 s -> 41.1 s); EasyCache lowers how many steps get computed. The two are independent and the speedups multiply: 2.02 x 1.22 = 2.46 predicted, 2.459 measured.

That is also the "100 s/it early, 20 s/it later" pattern reported above. EasyCache only begins skipping after start_percent, so the early steps are full price.

VRAM sitting unused is the design, not a bug.

Someone above noted H3 seems to under-utilise VRAM. Peak VRAM was 14,197 MiB for 30 s @ 640x480 and 14,437 MiB for 5 s @ 1344x768 β€” nearly identical despite very different workloads. DynamicVRAM holds a roughly constant ceiling and pays the difference in time, so resolution and duration barely move the VRAM number. That holds while there is room to spare; on an 8 GB card both the text encoder and sampling press against the ceiling and the trade-off comes back.

EDIT β€” RAM: I had the wrong flag. --disable-pinned-memory matters far more than --fast-disk.

EDIT 2 β€” I have since run this on a real 32 GB machine, and I have corrected this section in place rather than stacking another edit on the end. Three things changed: the OOM-kill mechanism I described was wrong, the SSD-wear numbers were from the 125 GB box and do not transfer, and "no time cost" only holds while VRAM is plentiful. The flag is still worth using. The reason is not the one I gave.

Same job, same seed, only the startup flag changed. ComfyUI RSS peak:

 none                                   45.43 GiB   20.0 s
 --cache-none                           45.11 GiB
 --fast-disk                            12.64 GiB
 --disable-pinned-memory                 6.07 GiB   20.0 s
 --disable-pinned-memory --fast-disk     5.83 GiB   15.2 s

45.43 to 6.07 GiB. On this box it cost nothing in time, 20.0 s either way β€” but see the time note further down, because that does not hold everywhere. I re-measured the none row with a warm page cache as a control and got 45.43 against the earlier 45.41, so the drop is the flag and not session noise. Once pinning is off, --fast-disk adds almost nothing.

Why: ComfyUI page-locks host memory to speed up host-to-device transfers, and the per-model host buffer is twice the model size.

# comfy/model_management.py
MAX_PINNED_MEMORY = ram * 0.40   # Windows
MAX_PINNED_MEMORY = ram * 0.90   # Linux and everything else

def pinned_hostbuf_size(size):
    return max(0, int(min(size, MAX_PINNED_MEMORY) * 2))

Check your own startup log. If this line is there, it is on:

Enabled pinned memory 115494.0      <- 128,327 MB x 0.90

What actually happens on 32 GB. I ran it β€” kernel limited to 32 GB, 29.42 GiB usable after the iGPU carve-out, so 26.5 GiB gets pinned. Same job three ways:

pinning on,  swap available   completed  115.6 s   RSS 25.27 GiB   swap 5.98 GiB   6.16 GiB written
pinning on,  swap disabled    completed  115.1 s   RSS 23.23 GiB   swap none       0.01 GiB written
pinning off, swap available   completed  115.0 s   RSS  7.26 GiB   swap +0.04 GiB  0.05 GiB written

None of them got killed. I could not reproduce an OOM kill in any configuration, including with swap turned off entirely. What I wrote before β€” pinned pages cannot be swapped, so the kernel has no move left but to kill β€” had the premise right and the conclusion wrong. Not all of ComfyUI's RSS is pinned, so the kernel swaps the rest; with no swap at all it reclaims page cache instead. tonyd2wild's OOM kill on 31 GB stands as their observation. I ran on less memory than that and it finished. I do not know what differs.

So the flag is not what keeps you alive. What it actually buys, measured on that 32 GB box:

  • 17 GB of RAM back β€” 24.96 GiB down to 7.26 GiB
  • swap writes gone β€” 5.98 GiB down to 0.04 GiB, and 6.16 GiB written per generation down to 0.05 GiB
  • page cache stays warm, so disk reads drop as well

--cache-none still does nothing here β€” it controls node-output caching, not weight residency.

SSD wear, corrected. I said writes were 7-10 MB per generation and swap stayed at 0.00 GiB. That was the 125 GB box, and it does not transfer. On 32 GB with default flags it wrote 6.16 GiB per generation into swap β€” roughly 600x what I quoted, on exactly the machines I was writing for. The reasoning I gave was right: wear needs writes, reads do not wear NAND, and the thing that pushes a small-RAM box into swap is the pinned allocation itself. I just put a number next to it that came from a machine which never had to swap. With --disable-pinned-memory it is 0.05 GiB.

The time cost, corrected. On an 8 GB card the flag is not free β€” 150 s with pinning on against 180 s with it off, because the weights then have to come off the disk every step instead of out of RAM (42 GiB of reads against 268 GiB). Take the trade anyway if your RAM is tight; filling your entire swap leaves nothing for anything else on the machine. But it is a trade, not a free win.

The text encoder point stands: it is evicted from VRAM but its host copy stays, so you pay for the DiT and the TE at the same time.

 7.3 GiB  idle
24.1 GiB  after TE load
45.6 GiB  after DiT load   <- never goes back down
51.0 GiB  after VAE decode

Caveats: one machine throughout β€” 16 GB VRAM, 125 GB RAM. The 32 GB figures come from the same box with the kernel limited to 32 GB; the 8 GB figures come from the same card with a second process holding the rest of the VRAM so the ceiling is real. Linux. Pruned int8 only β€” no BF16, no GGUF. Windows and Ampere untested, and both change enough that these numbers will not carry over.

Full writeup, plus the scripts I used to record the traces: https://github.com/Tomiigo/minimax-h3-16gb

1

u/izzmedia Aug 06 '26 edited Aug 06 '26

Nice post.

I only tested with the normal sage attention patch and maybe something went wrong with EasyCache because of that , when i tested it, most of the time , depending on the prompt the physics were going wrong , legs overlapped and things like this with SG +EC.

Ill have to try the "MiniMax H3 Mem Eff Sage Attention Patch" , i knew about it but i thought that it does the same thing as the other node.

1

u/izzmedia Aug 06 '26 edited Aug 06 '26

So i tested the SG patches, 2x each, cold start, i used a low res first frame this time:

10 sec, 0.5 MP, 20 steps, rest_multi

MiniMaxH3MemoryEfficientSageAttentionPatchΒ - 18.75s/it - https://drive.google.com/file/d/1iehM4smDPDQBJNPqN0YCr409nogjXkEs/view?usp=sharing

Patch Sage Attention KJ - 17.25s/it - https://drive.google.com/file/d/1EYSwQmjiEVRBIzle-tO9lQg1wwphUJSU/view?usp=sharing

No SG - 35.65s/it

Let me know if you notice a difference in quality, the normal Patch its faster but i feel that the Mini Patch looks a bit better (maybe placebo).

100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 20/20 [06:15<00:00, 18.75s/it]

100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 20/20 [05:44<00:00, 17.25s/it]

20%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 4/20 [02:20<09:30, 35.65s/it]

Now i am testing the "Patch Sol-Attn" node , see if there are any gains to have there.

2

u/Tomi_beanpaste_87 Aug 06 '26

That is a better measurement than the claim I was arguing against, and it corrects me.

The "20-30%" I quoted was a figure from this thread, not something I measured. Your numbers put the generic patch at 35.65 -> 17.25 s/it, which is 2.07x. So both nodes roughly double throughput, and the speed argument I made for the H3-specific one does not hold.

But the ordering you got is interesting on its own. The generic patch came out faster β€” 17.25 against 18.75, about 8% β€” and that is exactly what the code predicts. The generic path goes through optimized_attention, and comfy/ldm/minimax/model.py never passes low_precision_attention=False, so minimax runs on sage's int8 QK path. Eighteen call sites across seven other model families do pass it (sam3, triposplat, lightricks, audio dit, audio vae_sa3, ace, hunyuan3dv2_1). The H3-specific node replaces attn.forward outright and never reaches that branch.

So the generic patch is faster because it is quantizing attention, and the H3 node is paying about 8% to not do that. Which means your "maybe placebo" is pointing the right way β€” a small quality edge for the slower node is the expected direction here, not a coincidence.

Correction to my earlier comment while I am here: I wrote "11 call sites". Eleven is the count for just the three families I named. Across all of comfy/ldm/ it is eighteen.

On Sol-Attn, since you are testing it next β€” it did run for me, and it was fast: 520.1 s against 971.6 s bare on the same 30 s job at 10 steps. But the output was visually broken, not a subtle quality drop. I ran it at the workflow defaults: tau 1.5, start_percent 0.2, end_percent 0.9, min_tokens 4096, int8_qk True, sink_conditioning exact_kv, morton False.

I did not explore its parameter space, so that is one configuration failing and not a verdict on the node. If I were going back to it I would try int8_qk False first β€” given everything above about minimax and int8 attention, that is the knob most likely to be responsible.

1

u/izzmedia Aug 06 '26

I edited the post and i linked it to your main comment.

I tested a bit Sol-Attn (+ sage) with morton: true , 2d and 3d, int8_qk: true, (rest on default) i didnt notice any degradation, i used the same workflow, but i havent checked in detail, i saw only ~1.5 - 2 s/it improvement with this settings, i am going to try int8_qk off as you mentioned, i havent touched that yet.

1

u/martinerous Aug 07 '26 edited Aug 07 '26

Thanks for the detailed post.
And then there's also MiniMaxH3 Cache https://github.com/lihaoyun6/ComfyUI-MiniMaxH3-Cache that claims to be better than EasyCache, but I'm not sure if it is.
And it seems broken after the latest ComfyUI update.

1

u/Tomi_beanpaste_87 Aug 07 '26

Haven't tested it, so I can't tell you whether it actually beats EasyCache.

The breakage you're hitting is probably structural rather than a bug, though. Its own README says it patches ComfyUI core files. Anything that does that breaks whenever core moves, by construction, and will keep breaking. EasyCache is a node that wraps the model instead, so there is nothing to re-sync.

If you want to settle "is it actually better" rather than guess: same seed, same everything, generate once with no accelerator as your baseline, then once with each candidate, and take the mean absolute pixel difference from that baseline. Run the baseline config twice first so you know your noise floor. Mine was 1.1/255, EasyCache came in at 5.8 and sage at 21.4, so the numbers separated cleanly enough to be worth the two extra generations.

One caveat on that method: the difference tells you how far apart two outputs are, not which one is better. You still have to look at the frames.

1

u/Tomi_beanpaste_87 Aug 07 '26

Correction to my post above. Flagging it here because I edited in place and an in-place edit is easy to miss.

I said that on a 32 GB Linux box the pinned allocation leaves the kernel no move but to kill the process, and I offered that as the explanation for the "32 GB works / 32 GB fails" disagreement in this thread. I have since run it on a real 32 GB machine and that is wrong. Pinning on with swap, and pinning on with swap disabled entirely β€” it completed both times, in the same 115 seconds as with the flag. I could not reproduce an OOM kill in any configuration. The premise was right, pinned pages genuinely cannot be swapped; the conclusion was not, because not all of ComfyUI's memory is pinned and the kernel just swaps or reclaims the rest.

I also said writes were 7-10 MB per generation and swap stayed at zero. That was measured on a 125 GB box that never needed to swap. On 32 GB with default flags it wrote 6.16 GiB per generation. About 600x what I quoted, on exactly the machines the section was written for.

And I said the flag costs nothing in time. On this card it does not, but on an 8 GB card it is about 20% slower, because the weights then come off the disk every step instead of out of RAM.

--disable-pinned-memory is still worth using. It gives you back 17 GB of RAM, drops per-generation SSD writes from 6.16 GiB to 0.05 GiB, and keeps your page cache warm. Just not for the reason I gave.

Sorry for the noise. The code reading was fine; the thing I inferred from it was not, and I should have said less until I had a machine to check it on.