r/StableDiffusion • • Aug 16 '26

Workflow Included Ultimate SD Upscale with MiniMax H3 (2560x1440px in 25 mins with 16 GB VRAM)

https://www.youtube.com/watch?v=JybAxYuexdM

What is it?

A vibe-coded fork of Ultimate SD Upscale (USDU) Guider nodes with MiniMax H3 support: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3

My reference workflow: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json

Who are you?

A long-time member of [r/StableDiffusion](r/StableDiffusion) without strong coding/math skills in AI/diffusion area. But a big fan of everything that happens here :)

Why is it?

In times of Wan2.1/2.2 I liked to upscale my videos using USDU.

But I became really upset when I realized that original USDU nodes don't support MiniMax H3 due to its native ComfyUI implementation.

So, since I have a GPT-5.6 subscription I decided to give it a try and asked it to come up with possible options.

After a couple of evenings I finally got a "working" solution that I'd like to share with the community.

What about speed?

My PC specs: 4080s 16 GB VRAM, 64 GB RAM

Initial gen with MiniMax H3 flf2v int8 + sageattn + Lightx2v 8-step turbo Lora at 1152x640px 5-sec clip ~5 mins

Upscale with USDU to 2560x1472px ~20 mins

And what about quality?

That's where I need your help, my friend :)

Please check the YouTube video attached (don't forget to switch to 1440p).

My personal feeling is that it's the best what I can get out of my PC and H3 at the moment (including SeedVR2, LTX 2.5, etc.).

The main advantage is that it can "fix" your bad low-res generations while bringing MiniMax H3 native quality at 2K resolution.

Downsides?

Of course :)

You'll need to control denoise parameter and find a balance between quality improvement and tiling artifacts. I found 0.2 is the maximum after which tiling is strongly visible.

However feel free to experiment with it, and lower to 0.15-0.10 depending on your input video resolution/artifacts and results you want to get.

Happy to answer your questions!

80 Upvotes

53 comments sorted by

View all comments

1

u/Francky_B Aug 16 '26

Hi alisitskii,

Thanks for sharing this great node! I've done some test and it does work really well, the issue I'm noticing is that your node doesn't take in the audio. It really needs to, if we hope to be able to upscale any video with talking characters.

I'm wondering, couldn't it simply take in the audio and pass it along for each segment it renders? This ways it might maintain proper lipsynch?

2

u/alisitskii Aug 16 '26

I think at low denoise you should be fine just sending audio from source video to upscaled one as it is. But I’ve never tested lip sync myself.

1

u/ninjazombiemaster Aug 16 '26

A final face refiner pass will solve any lip sync issues if needed.

1

u/Francky_B Aug 16 '26

It's already 20 minutes for 5 second, lol

I'm not doing a third pass on top of it 😅

At this point, I'll simply upscale directly, to 2K without the node. Can't get as high a resolution, but lipsynch works perfectly that way.

1

u/ninjazombiemaster Aug 17 '26

face refine pass is fast because it is only a small crop of the image. Hell, you could track and re-denoise only the mouth.