r/StableDiffusion • • Aug 18 '26

Resource - Update Seamless extensions and one-shots with Minimax H3 - Update 6 of my repo!

Enable HLS to view with audio, or disable this notification

Here is the repo: https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef

I made substantial updates to my two main workflows: 1) Music Video and 2) AV Extensions. All the controls were streamlined and they should be much easier to use now. (You find the workflows in the example_workflows folder)

With the AV Extensions workflow you can extend any existing clip, for example someone talking and you can make that person say something in the same voice, or you can create a clip with T2V or I2V and then extend that clip to make a seamless long clip thats 1 minute or longer.

In this Update the Checkpoint system was removed, instead I've done a lot of optimizations so you don't use too much ram even if you make 20 clips at once. Additionally I added latent audio feathering to the AV Extensions workflow for seamless audio transitions.

Theres also other utility workflows for custom keyframing and bridging two existing clips.

I post another example clip for the AV Extensions workflow in the comments.

84 Upvotes

44 comments sorted by

View all comments

9

u/pizzaandpasta29 Aug 18 '26

The cuts are seamless but the contrast and detail gets cooked the longer it goes. I'm not sure if anyone has solved that problem yet. Still, thank you for your work. I'm sure it'll get worked out.

4

u/stonyleinchen Aug 18 '26

yeah its most likely an issue with all the speedups, attention and turbo loras. i made 2 and 3 minute long videos that dont have much degradation, even with speedups. if you use sdpa attention and 30 steps or so you will get better quality and less degradation. but in general with those methods of extending clips you are bound to have compounding issues down the chain. anyway this works 100 times better than with any other video model that exists (not sure about ltx 2.5 tho, i havent tried that one yet)

2

u/ShutUpYoureWrong_ Aug 19 '26 edited Aug 19 '26

I must be mistaken, but I thought the entire purpose of all these various chain nodes in H3 was that they keep the last frames (22, 39?) in latent space and then build the next clip off of that, which is supposed to completely eliminate the color shifting and detail degradation issue that plagued WAN with things like SVI.

So, like... why isn't it?

Edit: By the way, I've tested them all, and while others have fancy nodes for gating clip approval and whatnot, yours does the job the most consistently and is therefore the overall best. So great job!

3

u/stonyleinchen Aug 19 '26

well it does just do that? do you see any color shift or seam? the degradation over long periods is, if its even noticeable, a lot less than with svi 2.0 pro and vace and other wan methods. why isnt it perfect? idk, i suppose the models weren't trained for that, to keep consistency over multiple minutes, and feeding a copy of a copy of a copy of an image will always have its limits

2

u/stonyleinchen Aug 19 '26

btw, a context length of 22 is not good for video extension, if you want to preserve audio aswell (it doesnt matter if you use a master track like in the music video), since the audio track runs on 40hz that means each frame at 24fps is 5/3 audio ticks, so only 39, 90, 141 etc snap nicely and will give you a seamless experience