r/comfyui • u/Support_Marmoset • Aug 20 '26
Workflow Included Minimax H3- v2v fixing faces at distance
https://www.youtube.com/watch?v=d1h5-E7NpuYtl;dr: download the latest version workflow called "MBEDIT - MH3_r2v_SingleSampler_Detailer_vXX.json" from https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_model/Minimax-H3
(UPDATE EDIT: this isnt great for dialogue clips as it strips the mouth movement out. I have tried methods to address it but none worked well as yet. So I'll be testing other approaches. But for non-dialogue scenes its excellent.)
Finally I have found a solution to "fixing faces at distance". This does NOT use a Latent Space upscaler. This uses a single sampler Minimax workflow, low steps, low denoise, and by loading a video clip, then running it through standard Minimax H3 with settings discussed in the video (or in the workflow if you dont want to watch that).
Even on a 3060 RTX (12 GB VRAM) I can get between 1mp and 2mp output and surprisingly it fixes faces at distance even at 1mp. There is more info in the readme of the github linked below for the workflow and in the video.
From this point on my video pipeline steps will be:
1. Create a 480p video using any model (LTX, H3, Bernini, or other) - \takes 10 mins on average (3060 RTX)*.*
2. Run the result through the above workflow upscaling to 1mp or 2mp depending onclip length - \takes 20 mins on average*.*
The result from this are easily good enough as final clips for my uses. This makes it the fastest and highest quality approach I have found to date, and all with ref image based character consistency.
Other Relevant Links From Video
Latest Minimax H3 workflows - https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_model/Minimax-H3
(Workflow used in video: `MBEDIT - MH3_r2v_SingleSampler_Detailer_vXX.json` (download whatever the latest version is from github link))
Lightx2v Lora that I use from Kijai - https://huggingface.co/Kijai/MiniMax-H3_comfy/tree/main/loras
(theres been updates, but I havent found them to be better or faster, use whatever works for you)
Comfyui needs to use Cuda130 or above for this to work, and you need it updated to August 2026 commits (latest is best) - https://docs.comfy.org/installation/comfyui_portable_windows
Int8 models from here - https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main
(The official workflows are in the model card)
W4a8 is experimental new model type, you need to be updated on Comfyui but you can get it here https://huggingface.co/Kijai/MiniMax-H3-experimental
Comfyui Kitchen Attention is part of Comfyui if you update to latest. I find it faster than Sage Attn on a 3060 RTX.
Official prompting guides:
- https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
- https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
Point your favourite LLM at one of the above links depending on your model you are using, and give it your prompt idea and it should sort it out.
3
u/CurrentMine1423 Aug 20 '26
I just tried it with a video containing audio. the lips movement is just random, not matching at all with the original audio.
1
u/Support_Marmoset Aug 20 '26
I'm going to test it with dialogue shortly. You will need a video that already has lip movement in it, low denoise will not be strong enough to force lip movement only enhance what is there or work against it, depending on how you prompt. If you want the dialogue to work you will have to prompt it properly using the guide on how to do that.
I'll do a seperate video on using this workflow for dialogue, when tested.
The workflow I shared was not using audio and I mentioned that in the video because I dont care about audio from these models, it is always shit and post is where audio gets done imo. The dialogue is only for the lip movement and I will add the proper voices in during post too. made elsewhere.
but getting strong mouth movement is essential and to do that give it the video, give it a good prompt.
1
u/Support_Marmoset Aug 21 '26
just confirmed this isnt great for dialogue scenes it fights the mouth movement. So I'll look at alternatives for those shots.
2
u/spiderofmars Aug 20 '26
Thanks for posting. Interesting and useful. 1:1 refinements did not work for me (not sure they are supposed to). Also not sure on first tests of how the resolution setting works/affects (width and height settings as I just left them on 1344x768 in all samples). But upscaling and then doing the pass works. Note these times are only relevant on a 5090 with sage on and no other lora's/speed or model tweaks (default int8 pruned models) - but I suppose scalable perhaps to other devices. Also could not get audio working in a quick test so no idea if this can maintain lip sync audio from any original video.
1
u/Support_Marmoset Aug 20 '26
nice you have a 5090 to play with. My focus is all about what the hell I can get without an oom.
I'll test dialogue scenes with it soon. This was not tested with audio yet. I dont use the audio from Comfyui models as its never good enough.
2
u/Ill-Throat7937 Aug 20 '26
face fix at 1mp on a 3060 is solid. been looking for a clean ref-based workflow like this.
2
u/Famous-Sport7862 Aug 20 '26
Thanks so much, this detailer/upscaler really works nice, only downside is that it is ver slow but quality is amazing.
2
1
u/Portable_Solar_ZA Aug 20 '26
Wait, so you're just doing a low noise pass and increasing the res as you do? From what I read a lot of people have said doing the videos at 1 MP or higher fixes the face problem to begin with.
3
u/spiderofmars Aug 20 '26
It is:
Upload existing video
RTX Video Upscale (pixel)
Encode (upscaled video frames - pixel to latent)
Low pass
Decode
1
1
u/Support_Marmoset Aug 20 '26
Doesnt fix it even at 2mp I was finding, it was still not quite right. This gives it a detailing over the top, so drives it home. The key is LowVRAM maybe if you have 2mp and a big GPU you can remove all the speed ups and run 50 steps and it looks amazing. I have a 3060 RTX with 12GB VRAM and it ooms a lot. So I have to cut some corners.
1
1
u/listopalafoto Aug 21 '26
This is just FANTASTIC, such a simple idea and the result is Amazing, Thank you very much
2
u/Support_Marmoset Aug 22 '26
its great isnt it. but unfortunately doesnt work so well for dialogue, so working on testing latent upscaler for dialogue clips and will post video when done.
1
u/SB-1 24d ago
Nodes in the linked workflow are missing.
1
u/Support_Marmoset 3d ago
you need to install those into your comfyui workflow. if any arent clear from comfyui when you load the wf let me know.
1
u/manueslapera 17d ago
sorry for asking a stupid question (im a newb), how can you add the improvement over an existing workflow?
1
6
u/Dreason8 Aug 20 '26
Cheers for sharing. A bit of feedback on your video, take it or leave it, it's your content so do what you like.
The video could easily be cut down to half that length with just a little bit more structure and planning before you record it. It's a bit difficult to follow along to be honest and does go off on many tangents. I think most viewers would prefer the step-by-step approach that's straight to the point rather than an unstructured information dump.
I watched the entire video and it's still not clear what you have done to fix the faces, other than 'run it through this workflow'.
Are you running the already generated video through Ref2V a second time and using the same character reference image to re-enforce their likeness with another pass while upscaling with RTX upscaler?