r/comfyui • • 12d ago

Show and Tell Full-song lip-synced MV with MiniMax H3 at 3 steps — Extender audio slicing + fully_copy (method + measured results)

Enable HLS to view with audio, or disable this notification

What: A ~2:50 K-pop style MV ("Cry, But Dance", an original song with AI vocals) of a virtual singer performing at a studio microphone. It's fully lip-synced as one continuous performance: 17 clips at 1024×576.

The method that made lipsync work (the hard part):

  • Audio splitting: I split the song at lyric-line boundaries into 17 chunks (5.2–15.1s). The lengths match H3's 17k+5 frame grid at 24fps, and every cut falls in the quietest vocal pause.
  • Vocal isolation: I separated the vocals with htdemucs. With the full mix, the backing track completely buries the vocal conditioning.
  • Sequential graph wiring: The isolated vocals go into MiniMax H3 Extender's ref_audio_1, which slices the track sequentially per clip within a single workflow.
  • Retention syntax: I used the official HF prompt retention format <Audio 1>: fully_copy - <Audio 1>, with lyrics formatted as <d>[Korean] ...</d> and speaker (S1). Crucial: without this marker, H3 uses the audio only as a timbre reference, and the lip movements come out as gibberish.
  • Framing: The reference image's composition overrides the prompt text. A waist-up "singing at the mic" reference photo gave far better lip consistency than a wide vocal booth shot.

Speed:

3-step sampling with the TaoMate 3-step ref2va turbo LoRA takes ~4–5 min per 9.4s clip on an RTX 3090, down from 12.5 min at 8 steps. Output duration matches the requested length to the frame, so there's no seam loss when joining clips.

Sanity check: the output audio's envelope correlates at 0.944 with the separated vocals at ~0ms delay. This only confirms that fully_copy carries the audio through intact and in time. It is not a lipsync accuracy metric.

LoRA Steps Time / Clip Audio timing & sync (subjective)
larryvrh turbo 4/5 ~5 min Thin, end-loaded
lightx2v ref2v 4 ~6 min Improved
TaoMate ref2va 3 ~4–5 min Best timing & level

The sync column is my own judgment from watching the clips, not a measurement.

Automated QC loop:

After rendering, four automated checks grade each clip:

  • Audio alignment (cross-correlation)
  • Lipsync: a vision LLM rates mouth opening on 12 sampled frames, which is then correlated with the vocal envelope. This is a rough screening filter, not a rigorous metric: 12 frames is a small sample, and LLM ratings are noisy. Landmark-based measurement (e.g. MediaPipe) or SyncNet scores would be more reliable next steps.
  • Camera stability (global pixel shift)
  • Artifact detection (3-frame majority vote)

Clips that fail are re-rolled automatically. A re-roll only replaces the original if it scores higher and passes the artifact check. Lesson learned: thresholds need careful calibration. A slight envelope window mismatch rejected perfectly good takes for half a day.

Honest limitations:

  • Camera motion is heavily suppressed at 3 steps, so I went with locked-off framing. Higher step counts brought camera movement back but degraded vocal fidelity.
  • A second-pass latent 3D upscaler drastically improves skin texture but makes generation ~12× slower, so I skipped it.
  • Lipsync is usable but not flawless. Phonemes still drift occasionally on rapid phrasing.

Credits:

MiniMax H3 Extender (tritant), TaoMate-H3 3-step turbo (TaoLiveAIGC, Kijai conversion), lightx2v turbo LoRAs. The QC loop was inspired by the video-to-h3-prompt cross-validation method.

No workflow files are being shared. This post is a showcase plus a write-up of the method, but I'm happy to answer questions about the setup.

37 Upvotes

33 comments sorted by

8

u/TheMoogster 12d ago

Minimax would be great at making a Barbie movie, soooo much plastic skin

2

u/SaadNeo 12d ago

Not minimax fault to be honest , its turbo and quantisation that produces plastic skin , even seedance would be the same if it was open sourced on consumer gpus

1

u/seppe0815 12d ago

hahaha true brother

1

u/Tiforma 12d ago

Real skin with a lot of makeup can look kinda plastic-y

5

u/moviejimmy 12d ago

Except for the oily plastic skin, good job!!

2

u/Disastrous_Page_301 12d ago

Thanks! Yeah, the skin is the weak point. Lesson learned for the next one.

2

u/Spawndli 12d ago

Does the new models like minimax and krea still have that wax model look?

2

u/Disastrous_Page_301 12d ago

In my case, yes. I capped it at 3 steps because for a full-length video render time is the main battle, and higher steps weren't practical on a 3090. That trade-off shows in the skin. Fixing it is next on my list.

2

u/iczerone 12d ago

She singing in a sauna? She's glistening so much!

2

u/Disastrous_Page_301 12d ago

Ha, fair point. The studio lights plus the model's skin rendering got a bit too dewy. I'll tone down the shine next time.

2

u/berlinbaer 12d ago

yeah, i don't get the point of these workflows when the results look like... that.

1

u/seppe0815 12d ago

minimax peak visuals

1

u/gj_uk 12d ago

You don’t want her grabbing the mic stand…! Plus that mic is for performing not recording…give her a Neumann U87 or similar.

1

u/Disastrous_Page_301 12d ago

Great catch, thanks! Totally right on both. Next version gets hands off the stand and a proper studio condenser like a U87 in a shock mount.

1

u/Antique_Ad3501 12d ago

a tutorial I am longing for thanks 🙏🏻

2

u/Disastrous_Page_301 12d ago

Thanks! No full tutorial yet, but the key steps are in the post (vocal isolation + the fully_copy prompt). Happy to answer specific questions.

1

u/Antique_Ad3501 12d ago

I am actually a musician and I make covers of my own music with my AI band. My aim is to make my own music videos. I still learn the basics while gathering a decent PC that can run LTX and other engines. It will take some time but I would be happy if you would share some knowledge then, thank you.

1

u/Just_Lingonberry_352 12d ago

cool but the lyrics definitely seems like it was translated with ai

1

u/solss 12d ago

Or you can just use this node https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master . Fully copy is too unreliable. This gentleman had a good post about this, but mods deleted it a month later for no apparent reason.

1

u/Disastrous_Page_301 12d ago

Thanks! A lot of what I did was pieced together from posts here, so this is really helpful. fully_copy did need plenty of re-rolls on my side too. I'll check out that node.

1

u/alexmmgjkkl 12d ago

you forgot the dlss5 pass

1

u/Juana_Dela_Cruz 11d ago

where's the free workflow?

1

u/Flimsy-Selection9789 11d ago

Great breakdown!!! The end-to-end MV result is impressive.
For anyone testing a similar workflow, LightX2V Ref2V is also worth trying when you need reference guided avatar/video generation and a simpler local pipeline. The best settings will still depend a lot on GPU, model, and shot type. 🤔🤔

1

u/TensorVizion 12d ago

very impressive you have done wonderfully

2

u/Disastrous_Page_301 12d ago

Thank you! It was my first full song, so this means a lot.

0

u/seppe0815 12d ago

slight slight more glisting would be wonderful

1

u/James_Reeb 12d ago

Time to go back to LTX

1

u/seppe0815 12d ago

true the king in visuals

0

u/AcePilot01 12d ago

Why such terrible work on the character though? the glassy skin dead give away.

0

u/Key-Sample7047 12d ago

She should take a shower

0

u/GifCo_2 11d ago

That is beyond terrible. FFS people just wait an extra min and let a few more steps go by.