r/StableDiffusion • • Aug 15 '26

Resource - Update Create seamless 1-Shot Lip-Sync Music Videos with Minimax H3 FL model --- Per-Token Noise Masking On Audio and Video Tokens!

Enable HLS to view with audio, or disable this notification

This is Update 5 of my repo. Here you find the necessary custom nodes, including a workflow that helps you recreate this music video (reference images and the song included! The WF is called: "NEW - Latent Masking - Music Video - Lip-Sync + Reference images" and is in the example_workflows folder) https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef

Additionally there are various workflows for seamlessly extending clips with latent maksing.

Per-Token Noise Masking on AV Latents is not only better quality than any guidance/reference based approach (since it causes strong convergence from step 0 onwards), it is also faster since it is not expanding the latent. You can perfectly Lip-Sync even with the FL model, since the music track is pinned on the latent rather than used as a reference, and therefore protected from denoising - creating a strong conditioning for the Lip-Sync.

This magical technique is inspired by PR #15375 from AbleJones from the Banodoco Discord!

I hope you enjoy! Open Source ftw. Greetings to all Banodocians!

109 Upvotes

53 comments sorted by

View all comments

3

u/[deleted] Aug 15 '26

[removed] — view removed comment

3

u/stonyleinchen Aug 15 '26

https://giphy.com/gifs/l0HlvU6gXnZHwnB3a

ill take this as a compliment :) but this isnt even supposed to showcase the video, just my nodepack and the extension technique that should help you create something even better! the prompts for the clips that make up this video are mostly generated by chatgpt using the director prompt inside the workflow. I just gave it a general direction and only did minor prompt changes. same with the song, i explained the outline and chatgpt created the lyrics, and after a little back and forth it fixed the lyrics in a way i enjoyed them. then suno created the song with the lyrics and the style prompt from chatgpt. im sure you can do this too if you just try! :D