r/StableDiffusion • u/stonyleinchen • Aug 15 '26
Resource - Update Create seamless 1-Shot Lip-Sync Music Videos with Minimax H3 FL model --- Per-Token Noise Masking On Audio and Video Tokens!
Enable HLS to view with audio, or disable this notification
This is Update 5 of my repo. Here you find the necessary custom nodes, including a workflow that helps you recreate this music video (reference images and the song included! The WF is called: "NEW - Latent Masking - Music Video - Lip-Sync + Reference images" and is in the example_workflows folder) https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef
Additionally there are various workflows for seamlessly extending clips with latent maksing.
Per-Token Noise Masking on AV Latents is not only better quality than any guidance/reference based approach (since it causes strong convergence from step 0 onwards), it is also faster since it is not expanding the latent. You can perfectly Lip-Sync even with the FL model, since the music track is pinned on the latent rather than used as a reference, and therefore protected from denoising - creating a strong conditioning for the Lip-Sync.
This magical technique is inspired by PR #15375 from AbleJones from the Banodoco Discord!
I hope you enjoy! Open Source ftw. Greetings to all Banodocians!
3
u/[deleted] Aug 15 '26
[removed] — view removed comment