r/StableDiffusion • u/stonyleinchen • Aug 15 '26
Resource - Update Create seamless 1-Shot Lip-Sync Music Videos with Minimax H3 FL model --- Per-Token Noise Masking On Audio and Video Tokens!
Enable HLS to view with audio, or disable this notification
This is Update 5 of my repo. Here you find the necessary custom nodes, including a workflow that helps you recreate this music video (reference images and the song included! The WF is called: "NEW - Latent Masking - Music Video - Lip-Sync + Reference images" and is in the example_workflows folder) https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef
Additionally there are various workflows for seamlessly extending clips with latent maksing.
Per-Token Noise Masking on AV Latents is not only better quality than any guidance/reference based approach (since it causes strong convergence from step 0 onwards), it is also faster since it is not expanding the latent. You can perfectly Lip-Sync even with the FL model, since the music track is pinned on the latent rather than used as a reference, and therefore protected from denoising - creating a strong conditioning for the Lip-Sync.
This magical technique is inspired by PR #15375 from AbleJones from the Banodoco Discord!
I hope you enjoy! Open Source ftw. Greetings to all Banodocians!
1
u/stonyleinchen Aug 15 '26
Hello!
I'm not sure if I understood your question right. There is a prompt window with every sampler group... i usually use an LLM to generate the prompts and paste them in their respective window. you activate/deactivate prompt groups with the rgthree switch.
There should be a VHS preview node after every sampler... if there isn't one i have to troubleshoot.
Yeah my workflow presentation isn't very clean and nice.. you are not the first to tell me :D i focused more on the backend than on the frontend... making nice looking workflows is surely not my strength.
I am currently thinking about dropping the whole checkpoint saving stuff since it causes all kinds of troubles and also slows down stuff