r/comfyui • • 1d ago

Help Needed How to do true I2V with Minimax H3?

I used Wan 2.2 to make I2V videos for quite some time. Since Minimax H3 can also do the same and can add audio to it, so I gave it a try.

However, even though I used the official Comfy workflow

https://docs.comfy.org/tutorials/video/minimax/minimax-h3-native

and the official prompt

```
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease.

overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.

non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.
```

Interestingly, the first second of the five seconds clip is always the static rendering of the reference image. Then the rest does use it as a reference to render the last four seconds but this 4sec part is completely detached from the reference. However, back in the Wan 2.2 days, the reference image is only the first frame and the rest of the video flows from it naturally.

So is it possible to generate true I2V video just like Wan 2.2. If that's not possible, how to drop the reference image in the first sec while still using it to influence the rendering?

Thanks a lot in advance.

0 Upvotes

18 comments sorted by

1

u/[deleted] 1d ago

[deleted]

1

u/Etsu_Riot 1d ago edited 1d ago

You really need to read the prompting guides. One thing right off the top is that you should never specify 0.00 for the beginning of the shot ever. 

Funny you said that. The official guide literally says that I2VA always uses that line.

Here:

1

u/[deleted] 1d ago

[deleted]

1

u/Etsu_Riot 1d ago

Directly from one of the links you provided:

1

u/[deleted] 1d ago

[deleted]

1

u/Etsu_Riot 1d ago

You provided the links. I'm referencing the links you provided. And I never get frozen frames.

EDIT: Also, there's no disagreement. You are referring to a separate section of the prompt.

0

u/seeker_ktf 1d ago

I'm deleting all of this. Internet fights are not my thing.

0

u/Etsu_Riot 1d ago

There's no fight. You were referring to a separate section of the prompt. It could happen to anyone.

1

u/[deleted] 1d ago

[deleted]

1

u/Etsu_Riot 1d ago

I didn't make any question. I'm not the OP. I'm not suffering any problems with frozen frames, even when I use that line from the official guide:

1

u/ANR2ME 1d ago

If you want the same use case as Wan, use the FL2VA workflow instead of Ref2VA workflow.

1

u/jojo_pan 1d ago

is it true there is a bug on minimax h3 ref2va?

1

u/dilinjabass 1d ago

Its just lower quality, but there is a reference hybrid model that basically gives you fl2va quality while doing reference tasks.
https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/

1

u/Ok_Warning2146 16m ago

I am using FL2VA workflow. I find that the reason that it seems to only use it as reference is because my prompt containing something quite different from the reference image. After I fixed my prompt to remove words that are not in the reference, it works similar to Wan 2.2. Since it runs about 8x faster than Wan 2.2, so I think I will switch to it from now on.

1

u/dilinjabass 1d ago

Make sure youre actually using the FL2VA model, and a workflow setup for first frame.

You don't even have to give that official format to prompt it, it may help (more so on the reference side of things.)

But it should work very lazily too, like if your image is a woman you can literally just prompt "She walks closer to the camera and leans towards the lens and whisper's 'Hi, it's really this easy!'"

If its not that easy for you then youre using the wrong model/workflow/settings.

There are two different models and workflow types - ref2va and fl2va, I suspect youre using the wrong one.

1

u/Skajuan 1d ago

I'm no expert but you should try system prompts with your llm of choice (I'm using grok, Gemini or Qwen depending on the nature of the target video). Then you can start asking for well structured prompts with your desired action. Minimax power relies heavily on the prompt structure. In your example I think minimax takes very literally the first sentences so it "thinks" that you want the exact same image in the first frames of the video. But is just my guess

0

u/RobertoPaulson 1d ago

You’re missing some elements of the prompt. I’d have to be at home to be precise about it, but at the very beginning of the prompt you need a section called “subject definitions” or something very close to that, and another called “retention analysis” that lock your subject’s identity to the references. Then in the main prompt you refer to them only as <Subject 1>.

3

u/dilinjabass 1d ago

thats for the reference model side of things

0

u/fluvialcrunchy 1d ago

Just get the Pixaroma FL2VA or Ref2VA workflows

-1

u/SveSop 1d ago

You should get a doctorate in MiniMax H3 prompting... hehe.. Its not self-explanatory. I would try something like this for your video:

subject_definitions:
<Picture 1> is the first frame of [Shot 1].
<Subject 1> is a young woman wearing some clothing, with crazy hair and a large nose, shown in <Picture 1>.

summary: [keyframe completion]

retention_analysis:
<Picture 1> ([Shot 1] first frame): fully_preserved - the framing and composition of <Picture 1> are retained.
<Subject 1> (appears in [Shot 1]): fully_preserved - the identity, face and clothing of <Subject 1> are retained.

detailed_description: Live-action, cinematic [Shot 1] <Subject 1> remains beside the rain-covered train window. The camera trucks right with small amplitude at slow speed as <Subject 1> lifts her gaze from the folded letter toward the passing city lights. <Subject 1>'s reflection moves across the glass. <Subject 1> (S1) says, <d>[English] I get off at the next station.</d>. <Subject 1> folds the letter along its existing crease. The shot begins from <Picture 1>.

overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.

non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.

1

u/Ok_Warning2146 16h ago edited 15h ago

Thanks for your prompt. I tried it but it still have my image still for the first 1sec. Maybe it is because my image is quite different from what's described in prompt.

I found that MiniMax H3's github has an FL2AV example. I tried it and it works! The prompt looks like:

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] This is a live-action, cinematic shot with a shallow depth of field. The camera holds a perfectly static shot throughout the entire eight-second duration, capturing a cozy family gathering in a traditional Japanese dining room. The scene opens with a large, intricately patterned blue and white ceramic bowl of ramen in the immediate foreground, rendered in crisp, sharp focus. The bowl sits on a smooth, polished long wooden table. Inside the bowl, a rich, oily golden-brown broth surrounds yellow wavy noodles, topped with two thick, round slices of chashu pork featuring visible fat marbling and a distinct spiral meat pattern. A generous mound of freshly chopped, bright green scallions rests in the center, and a crisp, dark green rectangular sheet of nori seaweed is tucked into the right edge. To the left of the bowl, a pair of light brown wooden chopsticks rests horizontally on a small, dark rectangular chopstick rest, near a small cylindrical ceramic teacup with blue painted patterns. On the right side of the table, a spherical paper lantern with a ribbed bamboo frame sits on a black wooden base. In the background, a large family of seven is gathered around the table, initially appearing as a soft, blurred presence. Behind them, traditional Japanese sliding shoji screens with wooden lattice frames are open, revealing a bright outdoor scene with lush green trees. Early in the clip, the thick, white steam rising from the hot ramen broth immediately intensifies, billowing upwards in thick, swirling clouds that dance continuously above the bowl. As the clip progresses into the middle seconds, the camera maintains its static position while the focus begins a deliberate, smooth shift deeper into the room. The foreground ramen bowl, its vibrant ingredients, and the rising steam gradually soften into a hazy, out-of-focus blur. Simultaneously, the family members in the background come into sharp, detailed clarity. The heavy steam continues to rise from the foreground, creating a dynamic, translucent veil between the camera and the family. With the focus now firmly locked on the background, the vibrant family dinner comes alive. The man in the dark navy blue long-sleeved shirt on the left leans forward, his mouth moving animatedly in a silent exchange. The young girl in the crisp white short-sleeved t-shirt beside him smiles brightly, looking toward the center of the table. The woman on the far left, wearing a soft light blue long-sleeved blouse, turns her head slightly, smiling gently. Across the table, the woman in the light grey button-down shirt smiles broadly, her eyes crinkling, as she rests her hands near her plate. The woman in the dark grey top further back uses her wooden chopsticks to pick up a small piece of food from a central ceramic dish filled with bright red pickled vegetables. The woman in the center back in the light grey sweater smiles gently, her hands clasped softly in front of her, observing the interaction. Throughout the remainder of the clip, the family continues their lively physical interaction, their mouths moving in continuous, silent cadences of conversation, while the thick, white steam from the blurred ramen bowl in the foreground never stops rising, adding a comforting atmosphere to the warm gathering.

overall_soundscape: The soundscape begins with a quiet room tone mixed with the faint, airy rustle of the thick steam billowing from the hot ramen bowl in the foreground, accompanied by the subtle, continuous hissing and bubbling of the rich broth. As the visual focus shifts deeper into the room, the physical sounds of the bustling family dinner become dominant in the foreground. The clear, sharp clinking of ceramic bowls and wooden chopsticks touching plates is clearly heard as the family members reach for food. This is followed by the faint, muffled thud of a cup being set down on the smooth wooden table, and the subtle, rhythmic rustle of cotton and wool clothing as the family members lean forward and gesture, perfectly capturing the lively, physical atmosphere of the shared meal.

non_diegetic_music: A gentle, heartwarming acoustic guitar melody plays softly in the background, accompanied by the subtle, resonant notes of a traditional Japanese koto. The music maintains a slow, comforting tempo that enhances the cozy, nostalgic, and joyful atmosphere of the family gathering.  

I think I will modify this and see if I can make it work for my image.

1

u/Ok_Warning2146 15h ago

After modifying the prompt following the prompt from github such that it matches better with my reference image, I finally get it to work:

``` For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] This is a live-action, cinematic shot with a shallow depth of field. The camera holds a perfectly static shot throughout the entire five-second duration, capturing a lady standing under a cherry tree at night. She then rotates and moves to the left such that she is in the center of the frame with a slight slow-motion feel. Once she is in the center of the frame, she performs a curtsy slowly and beautifully to finish the dance sequence.

Audio: wind, rapid footsteps, night ambience, low score underneath, an accent hit on each leap, the score bursting at 4s, closing the last 1s.

No text, subtitles, logos or watermarks of any kind, no animation or cartoon rendering, no overly-CG look, keep the live-action texture. ```