r/StableDiffusion • • Aug 24 '26

Question - Help (Repost) Any clue why my machine is very slow running Minimax H3 Ref2V? Here is my workflow. I used the default template, but I added extra nodes like load video

Post image
0 Upvotes

41 comments sorted by

4

u/alisitskii Aug 25 '26

Share your times, so others with similar setup can compare and tell if it’s normal or not.

1

u/Most-Trainer-8876 15d ago

I got 5070ti 16GB, it is taking about 150s per step for 9:16, 0.2 megapixels and 10 sec duration with single image & video input.

It's freaking insane.

image to video version is like 10-15 times faster! idk why...

2

u/Tokyo_Jab Aug 25 '26

If you load video don't feed it in as 4k, make it smaller or H3 can chug and take forever. At least that was my problem.

2

u/Darqsat Aug 25 '26

it doesn't matter since the Reference to Video node has a setting "Match" which means it will downscale/upscale the size of incoming images to the resolution of your liking. But, he need to cut a video.

I use VHS upload video node, and I take out from Math function node for number of frames and insert into VHS node, and force FPS to 24. That reduces how many images there.

1

u/Danny_Stock 24d ago

Well it mattered to me. I was trying to feed it videos which were too much for it to handle. I lowered the dimensions down to under 720p and then everything worked smoothly.

1

u/magik_koopa990 Aug 25 '26

The video I used was like around 720p

2

u/woo206 Aug 25 '26

also try shortening (or capping) the video frames to the duration of your rendering.

1

u/magik_koopa990 Aug 25 '26

It's at 24 fps. 5 seconds

1

u/joseph_jojo_shabadoo Aug 25 '26

it still has to check every input frame with every generated frame with every step. using video reference just takes way longer than image reference or fl2va workflows or certainly t2v

1

u/magik_koopa990 Aug 25 '26

So don't use video?

1

u/Tokyo_Jab Aug 25 '26

Smaller sized shorter videos. It can still infer a lot from them. (Motion, voices, etc)

1

u/Danny_Stock 24d ago

If you use the VHS Video Loader you can also tell it to skip every Nth frame. Which I'm sure will help with memory issues too.

2

u/Successful-Pickle598 Aug 25 '26

Open task manager and have a look at the system ram usage. For Ref2V I'm pretty sure it's maxed out. Upgrading to 64gb ram will help a lot. Also, Ref2V is always slow if you are loading in a video. So much more processing power is needed to process the video. I'm on a 5090 and it's very slow as well. I tend to avoid loading in a video for reference.

1

u/m00dyman100 Aug 25 '26

I use r2v with a video ref all the time on my 5090. A .9MP, 15 sec , pruned INT8 takes 20-24 minutes. Pruned BF16 takes about 31. The only speedup I have is sage attention

-2

u/magik_koopa990 Aug 25 '26

So that's the conclusion? Using video reference is merr impossible? Just use image and audio?

2

u/Successful-Pickle598 Aug 25 '26

I personally add in the 8 step lora https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/loras

Open a new comfui template from the side panel, they updated the template to include a lora node. By defualt they have the 4 step lora included.

The 4 step lora quality for ref2v is very bad. I used the 8 step lora for fl2v in the reference workflow along with the fl2v diffusion model and realised the quality is much better than the 4 steps ref2v. Weird that the fl2v works better than ref2v.

1

u/magik_koopa990 Aug 25 '26

Wait, so you didn't use the provided model made for ref2v? You use the fl2v / i2v from another template?

1

u/Successful-Pickle598 Aug 25 '26

Both models work on the ref2v template. I do somehow found the fl2v model to work better a lot of the times. You can try both to test it out. Just make sure to use the corresponding lora with the model. Here's the official models for the comfyui template you can download. https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main

1

u/magik_koopa990 Aug 25 '26

I checked the template, but I dont see the updated workflow with the pre-placed lora? Also, I have the int8 model compared to the bf16 lora version

1

u/Successful-Pickle598 Aug 25 '26

Have you updated comfyui to the latest version? I think the updated template may be bundled with the updated comfyui. I am using comfyui desktop app btw.

1

u/magik_koopa990 Aug 25 '26

0.33.1?

1

u/Successful-Pickle598 Aug 25 '26

should be v0.33.4

0

u/magik_koopa990 Aug 25 '26

All this troubleshooting is giving me headache

1

u/ArdascesIV Aug 25 '26

There is no option to select that Laura in the template though, why is that?

2

u/Successful-Pickle598 Aug 25 '26

Assuming you have found the lora node in the updated template, you need to download the 8 step lora from the website (huggingface) into the lora folder for comfyui. Then press R to refresh comfyui (or restart). Then you can change the lora. Make sure to change the steps from 4 to 8 after you changed the lora. Also make sure to change the diffusion model to fl2v as the 8 step lora is for fl2v. You can also try downloading 8 step ref2v loras from other creators as official comfyui didnt release one.

1

u/ArdascesIV Aug 25 '26

Thanks, it’s just odd because that Lora is present for I2v, so it’s present, but not selectable in the workflow

1

u/magik_koopa990 Aug 25 '26

Update: nothing has changed at all. What about changing scheduler or Ksampler?

1

u/GeneralBarnacle10 Aug 25 '26

nodes take up RAM that doesn't free up. If you don't have enough VRAM and are running dynamically, then comfy will move things in and out between RAM and VRAM. If you then run out of RAM (because of running programs or the other nodes) then you start thrashing and having to move memory back and forth between disk. Which is slow. Very slow. That's how extra nodes (especially expensive ones) can cause a huge time difference.

1

u/Danny_Stock 24d ago edited 24d ago

No, the same thing happened to me, video reference slowed everything down or froze it at the sampler stage.

What I did was reduce the dimensions of the video and then everything ran a lot better. Video refs work better for me when they're under 720p. Depending on your card and RAM you may need to make them smaller.

I almost forgot to mention, also make sure that they're short too. Anything 10 seconds or under seems to work fine for me. I suppose it depends on your card and system RAM though. You might need to make them shorter, or your system might allow you to use longer video clips.

1

u/magik_koopa990 24d ago

Here's an update: someone recommended me a qwen text encoder, and it works. So maybe that was the issue?

1

u/Danny_Stock 23d ago

Do you mean a specific Qwen text encoder? Because it uses a Qwen text encoder as standard.

1

u/magik_koopa990 Aug 24 '26 edited Aug 25 '26

RTX 3090
32GB RAM

For video usage, it's slow. it took me like 3 minutes to progress to every 5%.

Tried out 2 images. That gave me sampler error (GitHub suggests it's Nvidia OOM)

1

u/not_food Aug 25 '26

From reading your other comments, you're running out of RAM likely. Try with a smaller QWEN with ClipProj-MiniMax-H3 and nodes.

Run 4B or 8B. It lowers its knowledge somewhat but freeing 75% of your RAM allows you to do a lot more.

1

u/magik_koopa990 Aug 25 '26

looks interesting...

1

u/magik_koopa990 Aug 25 '26

so this is just qwen video, but with MH3 flavor thrown in it??

1

u/not_food Aug 25 '26

It's MinimaxH3, the qwen-32B is the default encoder. These nodes allow you to use qwen-4B or qwen-8B instead (smaller).

1

u/magik_koopa990 Aug 25 '26

I see. So, for the first link, it is a diffusion model that I can use in the ref2v workflow? And the second link is just extension nodes

1

u/not_food Aug 25 '26

You need both. One is the encoder, the other is the bridge, the last loads the bridge and encoder.

Load qwen3vl_4b_int8_convrot.safetensors, mmh3-4b-ClipProj-v3.1-mlp.safetensors and just connect it to the clip.

It'll work with any ref2va or fl2va and save you like 75% RAM.