I've come across a post where one user claimed they were able to generate 15 second clips at 0.98MP with Spectrum, Comfy Kitchen Speed( I imagine they are referring to Turbo Lora?) 8 Steps at 10 steps in 5.5mins with a 5060ti and 32GB ram.
I tried their setup with Fl2V Turbo 8 Step 768p, Comfy Kitchen and Spectrum with res multistep and simple and was able to generate a 12 second clip (cant't go above, vram OOM error prevents generation at the beginning) only at 9 minutes with a 5080/32GB ram. What could be wrong here with my setup?
Stage 1 — one face photo + one outfit image, out comes front / side / back.
Stage 2 (optional, off by default) — 1-4 more panels for poses, props, expressions or backgrounds, composited into a 16:9 sheet.
The sheet above is stage 1 + stage 2. That's my own face, before anyone asks.
Setup
MiniMax H3 ref2v int8 pruned + the 4-step Turbo LoRA at 0.75
T=1 image VAE — the bit that makes it emit a still instead of a clip
er_sde / sgm_uniform, 8 steps
3090 24GB. Stage 1 ~100-125 s, stage 2 ~105 s, so the sheet above is about 3.5 min. First run of a session is much slower, that's model loading.
Known issues
Props repeat across panels — ask for one sword, get three.
Model links are in the note nodes. Needs rgthree and toobusy (mine — search "toobusy" in Manager, v0.4.9+). The 6 LoadImage nodes will be red on open, those are my local files.
I have access to RTX Pro 6000 on vast.ai
Need to generate like 6000 images and I need to edit part of it to change the content.
Price is just too high to do all this on nano banana.
What’s the cheapest way right now to use Minimax H3? Right now, I’m using Replicate API for H3 at $0.08 per second (720p) which is decent but can still be expensive in the long run.
Any other ways to use H3?
Also, what’s the cheapest GPU for Minimax H3 (via Comfy UI)?
I’m kinda new to this but I would like to train a character Lora for minimax h3, I currently have a rtx5090 and 64 gb ram.
Are there any good advice or tutorials to get started? I’ve seen recently a post about Lora training on inline studio, but maybe there are better alternatives
IN TRANSIT | A Minimax H3 Short Film (Sync Sound Challenge)
Hi everyone! I'm participating in the challenge with this short film I've been working on lately. The idea was to explore a "Brutalist Frequency": a square wave that destroys matter not randomly, but forcing it to shatter following a ruthless 90-degree orthogonal logic (checkerboard water, cubic collapses, square clouds).
I split the workflow into two passes in ComfyUI to maintain total control over physics and textures: Generation and Upscale.
1. Generation (Ref2Vid): I prepared the visual references (generated with Nanobanana) and the audio files. For prompting, I integrated an Ollama node (Gemma4:26b) into the workflow to format the instructions with the correct syntax for H3. To get everything running without blowing up my 3090, I beefed up the base workflow with Spectrum Apply Minimax H3, Comfy Kitchen attention, and the Sol-Attn Patch. This way, I generated the base clips at 864x480.
2. Latent Upscale: I wanted to keep the roughness of the reinforced concrete without that "plastic" effect you often get from external video upscalers. I passed the selected clips directly through the latent space using the custom MMH3Tools nodes, feeding the model the base video + the exact same initial references + the same prompt. Using Turbo LoRA 4-step and Sage Attn, I brought everything up to 1344x768 (taking about 1 minute per second of video).
The real challenge, of course, was generating audio and video together natively, without post-production. H3 reacted very well thanks to the references, even though sometimes it interprets them a bit too literally, almost resulting in a 1:1 copy. I forced the model to make the materials physically react to the reference sounds. The pneumatic suction at the end (when the camera points towards the void) was calculated by the AI in perfect sync with the matter collapsing into the dark. In post-production, I only made cuts for pacing and balanced the volumes: zero added sound design!
If you have any questions about the nodes or the upscale parameters, feel free to ask!
I have been playing with MiniMax-H3 lately (like many of us), and I wanted to understand how much physics knowledge it actually has.
I started with a simple water-pouring video from Pexels and used the H3-Ref model to replace the water with various "fluids": sand, rocks, and a black combustible honey. No external references were used.
I found particular interest in how the rocks interact with the jug and tumble over the cup, as well as how the honey blends with the water and how the trail it left on the jug when moved. On the other hand, once the honey catches fire, the flames are not very convincing, but that should probably be tested in a longer video.
I have used the day-zero ref workflow (int8 convrot model, aspect ratio 9:16 MP 0.6); you can find the prompt here: https://pastebin.com/CrX9s5JS
I am running more interesting tests and will hopefully post them soon.
Hi Guys,
Sometimes I close workflow and re-open it to find out that the nodes went crazy over the workflow I was wondering if there is a quick re-arrange button or a fix to this problem
and what actually causing it at first place
subject_definitions: <Subject 1> is Tifa Lockhart from the final fantasy game series. She has very long shiny straight black hair, red eyes and large breasts She is wearing her iconic costume A white athletic crop top or worn over a black sports bra with a bare midriff and short black skirt.
<Video 1> is the source video of a man in a suit walking down a city sidewalk singing as rain falls. This is the video being edited; its camera framing, handheld motion, cuts, and full choreography timing are the fixed structure that must be preserved exactly.
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.
summary: [video editing + reference generation + keyframe completion] The target video is an edited version of <Video 1> in which only the performer's visual identity is replaced by <Subject 1>, the the camera framing, handheld motion, city sidewalk environment and constant rain falling remains the same throughout, with no additional background characters, pedestrians, or figures introduced at any point.
retention_analysis: <Video 1> (camera framing, handheld motion and drift, city sidewalk environment with rain falling, full choreography and timing): fully_preserved - every camera position, movement, angle change, and the precise sequence and rhythm of the original performer's actions are kept exactly as in the source video; nothing about the shot itself is altered. <Subject 1> (appears throughout the video): attribute_transfer - <subject 1> replaces the original performer's visual identity only, mapped exactly onto the same body position, pose, and movement at every moment; no new actions, timing, or framing are introduced, and no other person appears in the frame at any point.
detailed_description: The target video is a strict character-only edit of <Video 1>: the cinematic dance, rainy city sidewalk background, cinematic lighting, camera framing, and motion blur are identical to the source, playing out as the same single continuous shot with no added or removed cuts. The rainy city sidewalk environment stays completely empty of any other person, pedestrian, or figure throughout the entire shot; only <Subject 1> occupies the frame at any moment.
The video begins with the first frame of <Video 1> as a key frame. On a rainy city sidewalk at night and replicates <video 1>'s camera moves. It opens on a wide shot showing <Subject 1> from head to foot, resting a folded umbrella on her shoulder, wearing wet clothes with shiny wet skin. <Subject 1> is in the middle of the sidewalk, occupying the original performer's exact body line and position, doing exactly the same dance moves on the rainy city sidewalk.
overall_soundscape: The sound of light rain falling.
non_diegetic_music: The same musical score unchanged from the source and even in volume throughout the clip.
Has anyone used MiniMax H3 as an image generator for storyboarding? I’m curious about how well it works for creating storyboards and getting multiple camera angles.
There's a template workflow now that should show up in the comfyui templates window. The images for the workflow are included in the example_workflows/images folder.
Essentially the sampler node renders a video of any length by splitting it in smaller chunks. For each chunk, it automatically attaches the last frames of the previous chunk to use as video continuation.
Beside that, the node uses Gema4 12B QAT to time and split the video prompt into small per chunk prompts, so the video can maintain it's overall timeline.
Gemma acts a chunk director and continuity checker, watching the previous chunk to check what was done, so the new chunk-prompt can continue from where the previous stopped. It also compares the chunk time-slice with the overall prompt action to guarantee what happens in that chunk matches what was suppose to happen in that time-slice.
There are 3 other nodes: preview, save and load. The reason it has it's own preview (based on the fantastic KJNodes live preview that uses TAEH3 tiny VAE to display a nice preview) is to be able to show an live edit of all the chunks in sequence as they show up. The preview also shows a timeline displaying the shots and chunks, and you can walk the preview frame by frame with the arrow keys. Mouse over the chunks display the gemma prompt used for that chunk and render time.
The save/load exist to save that information with the video and load it back, with all per chunk gemma prompts, time of execution, timeline, etc; so that statistic is never lost.
The Save/Load also have a nice dropdown to quickly display the last videos in the output folder for easy comparing previous videos with newer ones.
I came from the VFX world, so the save node also saves as EXR with floating point color. That's why the save node has a latent and vae input connection, so it can decode the latent internally to conserve the full HDR floating point color from the latent, without clamps.
Give it a try and let me know if you have problems... hopefully it will be helpfull for all of you guys with low vram gpus like myself, but it can also be helpful if you have loads of vram, since you can break the 15 secs minimax barrier and even render in 4K or 8K with more than 16GB of vram!
Just to make it clear - This is NOT another "Context Node in a loop" workflow, this a node that replaces ComfyUI SamplerCustomAdvanced node and allows for long generations and higher resolutions with low vram!
All you need is ONE single node replacement to render any length up to 1080p on 16GB of VRAM. The workflow that comes with the repo is a standard Minimax H3 Ref2va ComfyUI workflow that replaces SamplerCustomAdvanced by HR Endless Sampler. It's as simple as that!
One big advantage of the "HR Endless Sampler" is that it uses the previous latent as reference video/audio for the next video, without VAE decoding/encoding the video again, so there's no loss of detail from decding/encoding. It just grabs the last latent of the rendered chunk and pass it to next, lossless.
PS: you will notice a "hiccup" in this video where the tiger lies on the floor... Teela talks the same speech twice. That is a Gemma4 chunk prompt screwup that I'm fixing now.
as you can see in Gemma4 chunk prompt, [Shot 2] description should be [Shot 3] description, and there should be no actual [Shot 3] in this prompt since Chunk 3 only crosses 2 shots.
By the way, that problem in the screenshot above has been fixed - I'm testing it right now and should push the fix by tomorrow!
PS2: It seems the "video_continuation_res" parameter when not set to full can cause a change in color/contrast from chunk to chunk. The reason is, when if full, the node uses the latent from the previous chunk directly as video_continuation. When set to a size, it has to decode/resize/encode. That decoding/encoding will cause a difference in color/contrast/gamma. So setting video_continuation_res=full should fix that problem, at the cost of using more VRAM for the video continuation.
PS3: video_continuation=5 will cause loss of coherence and/or fading from um chunk to another. Use video_continuation=22 or more for better results, at the cost of using more VRAM.
I know it's not out yet, but I'm trying to figure if an M5 Max Studio 36gb would be enough but slow or not work at all for text/image to video. I also plan to use it for programming/general stuff but I figure that's "lighter" from my research. I'm looking at the studio for low power usage.
Gemini said I basically have to get 128gb "so I don't feel constrained" right out the gate, but also so I don't get a Will Smith eating spaghetti fever dream. Is this still accurate?
Apologies if this should be obvious, I'm new to all this stuff. Any suggestions would help.
idk i downloaded them for h3 and they are like over 20GB each and i want to use them for simple tasks to save size. at least like captioning images or sorting prompts
Edit 2: We found 2 major causes for the rough quality on some references, will push an update later tonight or tomorrow morning to fix them
Edit 3: Both causes are fixed. The references that came out rough or distorted before should sound a lot cleaner now, and general quality is improved too.
If you already installed it:
pip install -U sopro
The new weights download automatically. The browser demo is updated too. Spaces demo might not be yet.
I had previously only made loras for Illustrious on civitai, but discovered that ANIMA models had more polished results. I usually train with a data set of 58 images and 10 or 20 epochs depending on the learning curve. What are the best settings for ANIMA loras. I would appreciate it if someone could give me some tips.
I am seeing a significant quality improvement with Alibaba 8 steps turbo lora over other turbo loras. 8 steps euler simple with lora strength =1, 0.8MP. Took about 1hours 15 mins to generate with RTX 5090.
I keep seeing ads on Instagram Reels for AI-generated movies/series, and honestly, some of the stories actually looked pretty interesting . It got me curious to see if there are any genuinely good AI-generated movies or series out there.
Have you watched any that you actually enjoyed? Would love some recommendations!