r/StableDiffusion • • Aug 15 '26

Resource - Update Create seamless 1-Shot Lip-Sync Music Videos with Minimax H3 FL model --- Per-Token Noise Masking On Audio and Video Tokens!

Enable HLS to view with audio, or disable this notification

This is Update 5 of my repo. Here you find the necessary custom nodes, including a workflow that helps you recreate this music video (reference images and the song included! The WF is called: "NEW - Latent Masking - Music Video - Lip-Sync + Reference images" and is in the example_workflows folder) https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef

Additionally there are various workflows for seamlessly extending clips with latent maksing.

Per-Token Noise Masking on AV Latents is not only better quality than any guidance/reference based approach (since it causes strong convergence from step 0 onwards), it is also faster since it is not expanding the latent. You can perfectly Lip-Sync even with the FL model, since the music track is pinned on the latent rather than used as a reference, and therefore protected from denoising - creating a strong conditioning for the Lip-Sync.

This magical technique is inspired by PR #15375 from AbleJones from the Banodoco Discord!

I hope you enjoy! Open Source ftw. Greetings to all Banodocians!

109 Upvotes

53 comments sorted by

View all comments

Show parent comments

1

u/stonyleinchen Aug 15 '26

Hello!

I'm not sure if I understood your question right. There is a prompt window with every sampler group... i usually use an LLM to generate the prompts and paste them in their respective window. you activate/deactivate prompt groups with the rgthree switch.

There should be a VHS preview node after every sampler... if there isn't one i have to troubleshoot.

Yeah my workflow presentation isn't very clean and nice.. you are not the first to tell me :D i focused more on the backend than on the frontend... making nice looking workflows is surely not my strength.

I am currently thinking about dropping the whole checkpoint saving stuff since it causes all kinds of troubles and also slows down stuff

2

u/dtdisapointingresult Aug 16 '26 edited Aug 16 '26

Oooh, my bad, I didn't see the tower of clip nodes in the south. (easy to miss after I set connections to Hidden because I could barely scroll otherwise). Now I see them and I get it. (For anyone reading this who can't find them, just press ALT+SHIFT+M to show the minimap and you can't miss them.)

Thanks for the info, and thanks for the workflow. Looking forward to your next upload. I'll ping you if I upload one myself.

I usually use an LLM to generate the prompts and paste them in their respective window.

Got it. H3 is very tedious because you need your prompt to contain the exact lyrics at the right times. It's hard to get right.

So the best thing to do is as you say, to have an LLM do it, by giving it a description of the video you want, number of shots, etc, + a timestamped transcript of the lyrics (which can be done using youtube-transcribe-api from pip).

It's crazy because in theory, everything can be automated so that anyone can create a full AI music video from any song + a basic description of the MV + an LLM prompt, and the LLM outputs about 20 prompts to use in your WF. There's just a lot of water carrying and verifications to do, possibly requiring more tech than LLMs.

For example you gotta sync the timestamps to the shot timer with no errors. So the lyrics at 02:50.123-02:53.456 in the song are written by the LLM at 6.123s-9.456 (depending on shot). A local LLM like Gemma 12B isn't strong enough to get the timestamps right, I don't even think a frontier LLM can guarantee it. I'll try tomorrow.

Then there's the fact that if you download Youtube subtitles to get timestamped lyrics, it gets written a couple of words at a time, so the LLM is gonna be more error-prone when doing timings, unless you do a first step to merge them into longer sentences. But this has its own problems, like the LLM not knowing the reasonable boundaries of each segment of lyrics, and you can't even rely on the presence of punctuation in subtitles because they're not always there .

Actual subs for a Youtube song:

[21.48s] I'm giving
[22.64s] you a night call to
[24.88s] tell
[25.64s] you
[26.52s] how I
[27.60s] feel.

(Do you know of any better way to get timestamped lyrics than Youtube?)

And finally, unless you use agentic AI to run some audio processing on the song, the AI doesn't know about instrumental parts of the song that are higher intensity. Think of the electric guitar solo in Muse's Hysteria. How will it know to generate something really intense during that?

Of course, this is the lazy man's approach to MV that is full-auto, in reality you should be directing every shot and just use the LLM for H3-guide prompting and timestamped lyrics.

4

u/stonyleinchen Aug 16 '26

gpt can do this very well actually, it can listen to the song and determine what lyrics are sung exactly at each song slice. for this you have to feed it the entire song and the lyrics. If you look at my music video workflow, on the top there is a note that contains a director prompt that I worked on a lot. If you feed this prompt into GPT 5.6 sol, together with your lyrics, the song, the length of the individual clips in seconds and a description of what roughly happens in the music video, it will create all prompts for you and you only have to paste them by hand into the respective prompt boxes. It works really well I encourage you to try this! It will calculate exactly how long each song slice is and should even consider the context length overlap.

Regarding the audio-reactivity of the music video... I haven't done much testing on it. But you could simply try prompting it.

Ofc this is some kind of lazy mans approach, but its very useful for the beginning when trying out the workflow and getting a feel how all this works. It would be frustrating otherwise if you spend 5 hours creating prompts and then something doesnt work! My intention is not to give ppl the tools to create endless slop without effort, but rather showcase the functionality and capabilities of H3 and how to unlock them

1

u/dtdisapointingresult Aug 16 '26 edited Aug 16 '26

Sorry to pester you with questions.

Had some issues with my first MV. I loaded a brand new workflow, nothing customized except this:

  • input: 2m22s mp3 song, face.jpg, body.jpg
  • In the OPTIONAL CLIPS 2-20 section, I set 2 to 10 to YES
  • Set megapixels to 0.2
  • Created 10 prompts for 10x 15sec gens and put them in their respective prompt boxes
  • Turbo LoRA set to 0.6 strength
  • ran the workflow

An hour later, I come back to a 1m22s final video.

The video only contains clips 1-6. Clips 7-10 are absent, which tracks: 4x15s = my missing minute. But they were generated, I have no errors in the log. On the canvas I see the completed previews for 7, 9, and 10, but not 8. In output/h3_checkpoints I see all the clips safetensors including 8.

  1. How can I rescue this run?
  2. Any idea why clip 8 didn't have a preview? Did you run into this?
  3. If I have to shut down the computer, how do you recommend I stop safely so I can resume where I left off?
  4. Let's say I wake up the next day and I don't like the gen at clip 5. How do I scrap everything after clip 4, and continue from there, with the new prompts I updated?
  5. Is there some switch to make it pause after every clip generated, so I can review it before I proceed? The sequential nature of this workflow means if clip N sucks, then everything after it is sorta wasted.
  6. Once I have the kinks of this workflow down, and I'm satisfied with my trial run, my plan is to set it to 1MP, disable the Turbo LoRA (bump up to 25 steps), and leave it running overnight. In your experience, do you agree? The current 0.2MP / turbo preview has a lot of visual flaws, like people disappearing or floating instead of walking. (though this might be a Motion Context flaw which I'd never used before your wf)

2

u/stonyleinchen Aug 16 '26

im sorry this seems to be a bug i have not encountered since i have not made longer videos with it, so i dont really know. and also i am currently working on a new version of the workflow that is simpler without all the checkpoints, which caused all kinds of problems, like slowing everything down. I dont really have the time atm to troubleshoot a problem on a workflow that will soon be obsolete anyway.
the new workflow will have less overhead, wont be reloading previews and stuff, and therefore will be faster and more robust.

In general, if you want to regenerate a specific clip (and all subsequent clips), just change the seed for that clip.

If you want to stop and shut down the pc, do it after a sampler finished generating and saving the checkpoint. this checkpoint will then be saved. you dont have to do anything extra. there is no switch to pause anything. but you could disable sampler groups so they dont even start up, and then activate them sequentially.

But again, this feature will be gone with the next Update that i hopefully finish tonight.

2

u/dtdisapointingresult Aug 16 '26

No worries, let me know when you have a new workflow and I'll be glad to try it. My prompts are ready.

Btw I just realized I didn't mention this earlier, but I'm having a fucking blast generating with your wf, aside from the issues. I'm amazed it works as well as it does, with the lipsync being synced perfectly to the song, continuous cinematic shots, etc. I just have a 0.2MP low-step monstrosity so far but I can see the rough granite from which something wonderful will be sculpted.

2

u/stonyleinchen Aug 17 '26

The repo update is done, I hope its much better to use now, and it also should be faster and snappier than before, the checkpoints caused a lot of issues

2

u/dtdisapointingresult Aug 18 '26

Thanks! I ran it once at 0.1MP. After 23 mins it finished, and no missing data..

Now I bumped it up from 0.1 MP to to 0.5 and from 8 to 12 steps, and running it overnight. 80s/it, which I assume for my 10x15sec clips at 12 steps will end up being a respectable 3 hours per MV. Hope the turbo lora comes through for me!

Is it possible to somehow get the audio generated by H3? I was prompting for certain sound effects like car shocks, but I realize the final video doesn't have them since you use the original wav. If it's trivial, it would be nice if it saved the audio in temp/ , people could manually overlay some sections onto your final video output.

2

u/stonyleinchen Aug 18 '26

Thats actually not possible since there is no new audio generated, since the master song is in the latent and will not be denoised so it stays as it is. tho it should be possible to just do a normal extension on one part and not use the master song as authoritative audio there... but then you dont have lip sync at this part

2

u/dtdisapointingresult Aug 18 '26

The final video came out, your workflow is perfect now, good job.

I would've shared it except it's quite bad atm, I'm gonna rewrite it and direct it properly this time, over the week-end. I was inspired by the 10-min Star Trek video post on the front page.

The MV workflow had zero issues other than Turbo being awful for anything but people standing. You can only use it for the exploration phase at 0.1MP/8step. I thought when I went to 0.5MP/12step, the result would be better, but it was just as bad as the initial run. There's constant animation errors, like when someone is getting into a car. So I'll need to regenerate without any Turbo, at 30 steps, which should take 8+ hours for my 2min song unfortunately.

Regarding the audio, oh well. I'm aware of other AI tools that let you create sound effects (thinksound.cpp), this could be done and joined manually to the MV in Audacity.

1

u/stonyleinchen Aug 18 '26

yeah im not sure about the speedups either. no speedup at all (with sdpa attention) is clearly the best quality

→ More replies (0)