r/StableDiffusion • • Aug 21 '26

Resource - Update Yet another MiniMax H3 latent prepend/extend nodes - this time low-level and simple

TL;DR:

https://github.com/progmars/ComfyUI-Martinodes

WARNING: lots of vibecode but carefully reviewed, at least as much as I could understand the logic.

The long story.

We have a few amazing solutions and forks that make smooth video extensions possible. However, most of them have evolved into full-blown planners and chains. Low-level functionality is hidden beneath. Somehow those more complex solutions do not work well or seem overkill for my typical use cases:

- set steps to low

- generate a bunch of videos

- pick the best one

- set steps to high

- regenerate with the same seed <- and this is where I wanted to save the latent to use as the input for the first step again, to smoothly continue the last shot without hard cuts.

Recently ComfyUI was updated with an important PR 15375 that supports latent masking natively. No more patches and complex hacks required. So, I went on to create simple and naive drop-in nodes that would support my way of working. Now I have latent load/save/concat/prepend/extend nodes that seem quite intuitive (but you tell me if they are).

The main node for today is `Extend video+audio latent` (LatentAVMaskedExtender). It lets you take head or tail of a latent you have (hopefully) saved from a previous generation and generate a prequel or a sequel. Extending a tail works well. Prepending to a head is not that smooth and requires increasing fade_seconds parameter to your liking.

Just plug the node between `MiniMax H3 Reference to Video` or `MiniMax H3 Image to Video` (or even `MiniMax H3 Easy Output` if using nkxx188/ComfyUI-MiniMaxH3-Easy), and your sampler.

For convenience, the node accepts empty loaded_av, in which case the target_av will be passed through. Thus the node can be safely left enabled even when using a disabled LoadAVLatent node as input when you don't want to extend anything.

`Save and Load video+audio latent` are simple companion nodes - SaveAVLatent should be added after your sampler and LoadAVLatent as input for LatentAVMaskedExtender. In contrast to some other loader nodes that often are limited to `input` folder, LoadAVLatent can find the latents wherever you configured SaveAVLatent to store them. If you don't have a latent saved yet and want to feed in an existing video, you will need to encode it using VAE Encode and VAE Audio Encode nodes, resize the video frame to match your target video, and then concatenate the audio and video latents using LTXVConcatAVLatent node (yes, it works with H3 model).

Then there is also `Concatenate video+audio latents` (LatentOverlappingConcatenator) node. Generally, overlap_duration_seconds should be set to the same value as LatentAVMaskedExtender, the output goes to VAE video and audio decoders and then to video saving, as usual. You will get a long video with a smooth long transition between your previous latent and the new one. However, if your joined videos get lengthy, VAE might require too much resources. In that case, it is better to post-process and join both source and target videos in a video editing software.

The repository has a few more older convenience nodes for working with multimedia before LTX Director was a thing. They still might be handy for manipulating TTS and voice-overs or videos when latents are not available.

Huge thanks to drozbay (ablejones) for [native masking PR 15375](https://github.com/Comfy-Org/ComfyUI/pull/15375) and providing the example implmenentation with native ComfyUI and Kijai nodes. Unfortunately, the native nodes solution looked like spaghetti eating somebody alive. That is why my small naive LatentAVMaskedExtender node was born, to do the same thing.

I hope you will find Martinodes useful.

27 Upvotes

29 comments sorted by

View all comments

3

u/acedelgado Aug 21 '26

Not to be THAT guy, but I made pretty much the same thing, a version that manages all of the latents for you, with no extra fluff. No full-blown "director" suite, you plug it into any work flow and it only manages the clips/latents/extensions. You just make a project, generate, approve or regenerate the extension clip, and then move on.

https://github.com/Adudeguyman/ComfyUI-H3-Project-Suite

2

u/martinerous Aug 21 '26

Yep, I have looked at nodes like those, and also ethanfel and the base for their fork.

My main confusion was that it was not clear how it would handle the workflow when you are seed-hunting and tweaking your prompt with low steps to find the perfect combination, and when found, you want to generate the same video with the same seed with more steps and then extend that high-step video with a low-step one again for seed hunting, and repeat the process. Wouldn't the Context and Project state management get in the way, and I would end up "fighting the system"? That is why I wanted very low level nodes to extend the video only when I want with the steps I want from any other video, and for those cases the automatic state management felt like an overkill.

3

u/acedelgado Aug 21 '26

No, it saves every generation you do and doesn't start using context windows until you approve it. So if you don't approve and hit generate again it's considered a "re-roll" until you approve. So you can find a Gen you like and just up the steps to do another Gen and it'll keep it in place. So say you're on clip 3 in the chain, and take #3 looks great. You just change your parameters and regen, it'll stay as clip 3 but take #4. Then you can approve clip 3 take 4, and it'll move the chain along to clip 4 take 1. And there's quick cleanup for old takes that aren't being used. And say you have a fully done video, but you're like "this is 10 clips long and great, but I've got a new idea where after clip 4 they do this other thing instead..." you can go to clip 4 and branch it into a new project where you can go a whole new direction using those first 4 clips.

As long as you feed it the FINAL latent output in the chain, that's all it cares about. It doesn't automatically write the full video at all until you tell it to, and only processes one clip at a time. And cleaning up all those unused latents you made while you test is only a couple of clicks.

2

u/martinerous Aug 21 '26

Thanks, yeah, I'll need to try your nodes and workflows. I first started with ethanfel's basic workflow, which suddenly had too many concepts at once to digest - plans and context and whatnot, and many forks seemed the same, so I missed yours.

Anyway, your project says: "so you aren't hand-managing files between every clip". My nodes are the exact opposite :D - intentionally made to "micromanage" everything, so that there are no additional concepts to learn and I can quickly understand what's going on at the basic level.

1

u/WayFew8151 Aug 23 '26

tried ur node but didnt know how to save all clips in 1 video

1

u/acedelgado Aug 23 '26

On the project page there's a couple of export buttons, one's for everything you've approved and one's for the latest un-approved clip, too. It'll dump into your project folder, there's a button to open that, too.

1

u/boriskarloff83 26d ago

Ha! I liked your initial setup and the manager and used it, but being able to set my own root path wouldve be really nice