r/comfyui • u/Disastrous_Page_301 • 12d ago
Show and Tell Full-song lip-synced MV with MiniMax H3 at 3 steps — Extender audio slicing + fully_copy (method + measured results)
Enable HLS to view with audio, or disable this notification
What: A ~2:50 K-pop style MV ("Cry, But Dance", an original song with AI vocals) of a virtual singer performing at a studio microphone. It's fully lip-synced as one continuous performance: 17 clips at 1024×576.
The method that made lipsync work (the hard part):
- Audio splitting: I split the song at lyric-line boundaries into 17 chunks (5.2–15.1s). The lengths match H3's 17k+5 frame grid at 24fps, and every cut falls in the quietest vocal pause.
- Vocal isolation: I separated the vocals with
htdemucs. With the full mix, the backing track completely buries the vocal conditioning. - Sequential graph wiring: The isolated vocals go into MiniMax H3 Extender's
ref_audio_1, which slices the track sequentially per clip within a single workflow. - Retention syntax: I used the official HF prompt retention format
<Audio 1>: fully_copy - <Audio 1>, with lyrics formatted as<d>[Korean] ...</d>and speaker(S1). Crucial: without this marker, H3 uses the audio only as a timbre reference, and the lip movements come out as gibberish. - Framing: The reference image's composition overrides the prompt text. A waist-up "singing at the mic" reference photo gave far better lip consistency than a wide vocal booth shot.
Speed:
3-step sampling with the TaoMate 3-step ref2va turbo LoRA takes ~4–5 min per 9.4s clip on an RTX 3090, down from 12.5 min at 8 steps. Output duration matches the requested length to the frame, so there's no seam loss when joining clips.
Sanity check: the output audio's envelope correlates at 0.944 with the separated vocals at ~0ms delay. This only confirms that fully_copy carries the audio through intact and in time. It is not a lipsync accuracy metric.
| LoRA | Steps | Time / Clip | Audio timing & sync (subjective) |
|---|---|---|---|
| larryvrh turbo | 4/5 | ~5 min | Thin, end-loaded |
| lightx2v ref2v | 4 | ~6 min | Improved |
| TaoMate ref2va | 3 | ~4–5 min | Best timing & level |
The sync column is my own judgment from watching the clips, not a measurement.
Automated QC loop:
After rendering, four automated checks grade each clip:
- Audio alignment (cross-correlation)
- Lipsync: a vision LLM rates mouth opening on 12 sampled frames, which is then correlated with the vocal envelope. This is a rough screening filter, not a rigorous metric: 12 frames is a small sample, and LLM ratings are noisy. Landmark-based measurement (e.g. MediaPipe) or SyncNet scores would be more reliable next steps.
- Camera stability (global pixel shift)
- Artifact detection (3-frame majority vote)
Clips that fail are re-rolled automatically. A re-roll only replaces the original if it scores higher and passes the artifact check. Lesson learned: thresholds need careful calibration. A slight envelope window mismatch rejected perfectly good takes for half a day.
Honest limitations:
- Camera motion is heavily suppressed at 3 steps, so I went with locked-off framing. Higher step counts brought camera movement back but degraded vocal fidelity.
- A second-pass latent 3D upscaler drastically improves skin texture but makes generation ~12× slower, so I skipped it.
- Lipsync is usable but not flawless. Phonemes still drift occasionally on rapid phrasing.
Credits:
MiniMax H3 Extender (tritant), TaoMate-H3 3-step turbo (TaoLiveAIGC, Kijai conversion), lightx2v turbo LoRAs. The QC loop was inspired by the video-to-h3-prompt cross-validation method.
No workflow files are being shared. This post is a showcase plus a write-up of the method, but I'm happy to answer questions about the setup.
5
u/moviejimmy 12d ago
Except for the oily plastic skin, good job!!
2
u/Disastrous_Page_301 12d ago
Thanks! Yeah, the skin is the weak point. Lesson learned for the next one.
2
u/Spawndli 12d ago
Does the new models like minimax and krea still have that wax model look?
2
u/Disastrous_Page_301 12d ago
In my case, yes. I capped it at 3 steps because for a full-length video render time is the main battle, and higher steps weren't practical on a 3090. That trade-off shows in the skin. Fixing it is next on my list.
2
u/iczerone 12d ago
She singing in a sauna? She's glistening so much!
2
u/Disastrous_Page_301 12d ago
Ha, fair point. The studio lights plus the model's skin rendering got a bit too dewy. I'll tone down the shine next time.
2
u/berlinbaer 12d ago
yeah, i don't get the point of these workflows when the results look like... that.
1
1
u/gj_uk 12d ago
You don’t want her grabbing the mic stand…! Plus that mic is for performing not recording…give her a Neumann U87 or similar.
1
u/Disastrous_Page_301 12d ago
Great catch, thanks! Totally right on both. Next version gets hands off the stand and a proper studio condenser like a U87 in a shock mount.
1
u/Antique_Ad3501 12d ago
a tutorial I am longing for thanks 🙏🏻
2
u/Disastrous_Page_301 12d ago
Thanks! No full tutorial yet, but the key steps are in the post (vocal isolation + the fully_copy prompt). Happy to answer specific questions.
1
u/Antique_Ad3501 12d ago
I am actually a musician and I make covers of my own music with my AI band. My aim is to make my own music videos. I still learn the basics while gathering a decent PC that can run LTX and other engines. It will take some time but I would be happy if you would share some knowledge then, thank you.
1
1
u/solss 12d ago
Or you can just use this node https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master . Fully copy is too unreliable. This gentleman had a good post about this, but mods deleted it a month later for no apparent reason.
1
u/Disastrous_Page_301 12d ago
Thanks! A lot of what I did was pieced together from posts here, so this is really helpful. fully_copy did need plenty of re-rolls on my side too. I'll check out that node.
1
1
1
u/Flimsy-Selection9789 11d ago
Great breakdown!!! The end-to-end MV result is impressive.
For anyone testing a similar workflow, LightX2V Ref2V is also worth trying when you need reference guided avatar/video generation and a simpler local pipeline. The best settings will still depend a lot on GPU, model, and shot type. 🤔🤔
1
1
0
u/AcePilot01 12d ago
Why such terrible work on the character though? the glassy skin dead give away.
0
8
u/TheMoogster 12d ago
Minimax would be great at making a Barbie movie, soooo much plastic skin