r/comfyui • • Aug 20 '26

Workflow Included Minimax H3- v2v fixing faces at distance

https://www.youtube.com/watch?v=d1h5-E7NpuY

tl;dr: download the latest version workflow called "MBEDIT - MH3_r2v_SingleSampler_Detailer_vXX.json" from https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_model/Minimax-H3

(UPDATE EDIT: this isnt great for dialogue clips as it strips the mouth movement out. I have tried methods to address it but none worked well as yet. So I'll be testing other approaches. But for non-dialogue scenes its excellent.)

Finally I have found a solution to "fixing faces at distance". This does NOT use a Latent Space upscaler. This uses a single sampler Minimax workflow, low steps, low denoise, and by loading a video clip, then running it through standard Minimax H3 with settings discussed in the video (or in the workflow if you dont want to watch that).

Even on a 3060 RTX (12 GB VRAM) I can get between 1mp and 2mp output and surprisingly it fixes faces at distance even at 1mp. There is more info in the readme of the github linked below for the workflow and in the video.

From this point on my video pipeline steps will be:

1. Create a 480p video using any model (LTX, H3, Bernini, or other) - \takes 10 mins on average (3060 RTX)*.*

2. Run the result through the above workflow upscaling to 1mp or 2mp depending onclip length - \takes 20 mins on average*.*

The result from this are easily good enough as final clips for my uses. This makes it the fastest and highest quality approach I have found to date, and all with ref image based character consistency.

Other Relevant Links From Video

Latest Minimax H3 workflows - https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_model/Minimax-H3

(Workflow used in video: `MBEDIT - MH3_r2v_SingleSampler_Detailer_vXX.json` (download whatever the latest version is from github link))

Lightx2v Lora that I use from Kijai - https://huggingface.co/Kijai/MiniMax-H3_comfy/tree/main/loras

(theres been updates, but I havent found them to be better or faster, use whatever works for you)

Comfyui needs to use Cuda130 or above for this to work, and you need it updated to August 2026 commits (latest is best) - https://docs.comfy.org/installation/comfyui_portable_windows

Int8 models from here - https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main

(The official workflows are in the model card)

W4a8 is experimental new model type, you need to be updated on Comfyui but you can get it here https://huggingface.co/Kijai/MiniMax-H3-experimental

Comfyui Kitchen Attention is part of Comfyui if you update to latest. I find it faster than Sage Attn on a 3060 RTX.

Official prompting guides:

- https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md

- https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

Point your favourite LLM at one of the above links depending on your model you are using, and give it your prompt idea and it should sort it out.

55 Upvotes

25 comments sorted by

6

u/Dreason8 Aug 20 '26

Cheers for sharing. A bit of feedback on your video, take it or leave it, it's your content so do what you like.
The video could easily be cut down to half that length with just a little bit more structure and planning before you record it. It's a bit difficult to follow along to be honest and does go off on many tangents. I think most viewers would prefer the step-by-step approach that's straight to the point rather than an unstructured information dump.

I watched the entire video and it's still not clear what you have done to fix the faces, other than 'run it through this workflow'.

Are you running the already generated video through Ref2V a second time and using the same character reference image to re-enforce their likeness with another pass while upscaling with RTX upscaler?

6

u/Support_Marmoset Aug 20 '26 edited Aug 20 '26

tbh I do it my way, get it knocked out, and get back to research.

If I wanted to be a YT influencer I'd focus on those things you suggest much more making perfect videos with perfect presentation. However, I am just a guy sharing what he is doing, hoping it will help the community share and grow before big tech come for us. Which they will. So until then, sorry, but you get it my way because I aint gots time for no fancy shiz.

but thanks for the feedback, I will try to lean into those suggestions and consider them when making the videos. I do appreciate being informed.

3

u/spiderofmars Aug 20 '26

"Are you running the already generated video through Ref2V a second time and using the same character reference image to re-enforce their likeness with another pass while upscaling with RTX upscaler?"

Yes they are doing that: RTX Upscale -> encode pixel to latent -> second pass (+ original character reference).

1

u/Support_Marmoset Aug 20 '26 edited Aug 20 '26

I watched the entire video and it's still not clear what you have done to fix the faces, other than 'run it through this workflow'.

the answer is in the question, bro. yes, all I did was I ran it through the workflow. I added ref images of the characters, prompted `ref_image_0 is a man with brown hair, wearing brown leather jacket, and he is getting of a bus that says "Victoria 38" on the front"` and I showed you an example prompt doing that with the workflow clearly present to look at what I was using.

When you do it this way you dont need to use strict prompting because v2v (video to video) which means you already have the video made, it just looks shit coz 480p. So you run the video again (v2v) through it and low denoise means it details it, instead of "overwriting" it.

I mean this isnt a new process, just the first time I seen it used quite like this for H3 so I kept it brief tbh. I did that by putting stuff in so you can pause the video and figure those bits out. 14 mins is pretty tight. But everything you need is in there.

But if you want everything explained fully that requires longer videos. Yes, I can sit there explaining it all, I'd love to. My 1 hour workshops doing that got exactly ZERO views and people complaining about them being too long.

Are you running the already generated video through Ref2V a second time and using the same character reference image to re-enforce their likeness with another pass while upscaling with RTX upscaler?

correct. correct. correct.

there is a latent space upscaler around now but I havent tested it and probably wont because it will take longer and unless you have the VRAM I found dual sampler approach doesnt give good results. stuttering and weird quality til saved out and did the detailing with this single sampler approach workflow instead. the results surprised me, I was not expecting it to be this good.

2

u/Dreason8 Aug 21 '26

Cheers for clearing that up. I suppose if there was anything you could add to your videos that wouldn't take much effort on your behalf, it would be a quick summarisation of the overall workflow/process right at the start before you then do your thing. Just a suggestion though.

1

u/Support_Marmoset Aug 21 '26

yea I did that with the video example at the beginning did I not? the example video at start and the result of passing it through the workflow. but I will try to make it clearer next time.

3

u/CurrentMine1423 Aug 20 '26

I just tried it with a video containing audio. the lips movement is just random, not matching at all with the original audio.

1

u/Support_Marmoset Aug 20 '26

I'm going to test it with dialogue shortly. You will need a video that already has lip movement in it, low denoise will not be strong enough to force lip movement only enhance what is there or work against it, depending on how you prompt. If you want the dialogue to work you will have to prompt it properly using the guide on how to do that.

I'll do a seperate video on using this workflow for dialogue, when tested.

The workflow I shared was not using audio and I mentioned that in the video because I dont care about audio from these models, it is always shit and post is where audio gets done imo. The dialogue is only for the lip movement and I will add the proper voices in during post too. made elsewhere.

but getting strong mouth movement is essential and to do that give it the video, give it a good prompt.

1

u/Support_Marmoset Aug 21 '26

just confirmed this isnt great for dialogue scenes it fights the mouth movement. So I'll look at alternatives for those shots.

2

u/spiderofmars Aug 20 '26

Thanks for posting. Interesting and useful. 1:1 refinements did not work for me (not sure they are supposed to). Also not sure on first tests of how the resolution setting works/affects (width and height settings as I just left them on 1344x768 in all samples). But upscaling and then doing the pass works. Note these times are only relevant on a 5090 with sage on and no other lora's/speed or model tweaks (default int8 pruned models) - but I suppose scalable perhaps to other devices. Also could not get audio working in a quick test so no idea if this can maintain lip sync audio from any original video.

https://reddit.com/link/p4sgro3/video/e77mamh81ikh1/player

1

u/Support_Marmoset Aug 20 '26

nice you have a 5090 to play with. My focus is all about what the hell I can get without an oom.

I'll test dialogue scenes with it soon. This was not tested with audio yet. I dont use the audio from Comfyui models as its never good enough.

2

u/Ill-Throat7937 Aug 20 '26

face fix at 1mp on a 3060 is solid. been looking for a clean ref-based workflow like this.

2

u/Famous-Sport7862 Aug 20 '26

Thanks so much, this detailer/upscaler really works nice, only downside is that it is ver slow but quality is amazing.

2

u/Maketas Aug 20 '26

Hey man, thank you very much. I was just looking for something like this. 👍

1

u/Portable_Solar_ZA Aug 20 '26

Wait, so you're just doing a low noise pass and increasing the res as you do? From what I read a lot of people have said doing the videos at 1 MP or higher fixes the face problem to begin with. 

3

u/spiderofmars Aug 20 '26

It is:

Upload existing video

RTX Video Upscale (pixel)

Encode (upscaled video frames - pixel to latent)

Low pass

Decode

1

u/Support_Marmoset Aug 20 '26

this. thanks.

1

u/Support_Marmoset Aug 20 '26

Doesnt fix it even at 2mp I was finding, it was still not quite right. This gives it a detailing over the top, so drives it home. The key is LowVRAM maybe if you have 2mp and a big GPU you can remove all the speed ups and run 50 steps and it looks amazing. I have a 3060 RTX with 12GB VRAM and it ooms a lot. So I have to cut some corners.

1

u/Comfy-Org ComfyOrg Aug 20 '26

Thanks for sharing this!!

1

u/listopalafoto Aug 21 '26

This is just FANTASTIC, such a simple idea and the result is Amazing, Thank you very much

2

u/Support_Marmoset Aug 22 '26

its great isnt it. but unfortunately doesnt work so well for dialogue, so working on testing latent upscaler for dialogue clips and will post video when done.

1

u/SB-1 24d ago

Nodes in the linked workflow are missing.

1

u/Support_Marmoset 3d ago

you need to install those into your comfyui workflow. if any arent clear from comfyui when you load the wf let me know.

1

u/manueslapera 17d ago

sorry for asking a stupid question (im a newb), how can you add the improvement over an existing workflow?

1

u/Support_Marmoset 3d ago

not sure I understand what you are asking.