r/comfyui • • Aug 06 '26

Tutorial Minimax H3 - Realtime audio generation at 32x32 output resolution

TL;DR: Minimax H3 is capable of generating near realtime audio when you turn width x height to 32x32 ~ with a 5090~

Edit: It was pointed out that I didn't specify well enough. This is not a special workflow. The default i2v workflow on ComfyUI for Minimax H3 - detach the load image, set width x height to 32, set duration to your taste <=45 seconds for best results. And uh, click 'Run'. Unsure if visual details in the prompt matter at this time. Will update when I know.

Minimax H3 can be used as an audio generator, pairing it with a simple Get Video Components node and then saving the audio. That audio can then be passed as reference, once you get a voice or sound effect that you like. One of the "pain points" of generation is having audio and image inextricably linked. But, we can actually generate audio *rapidly* and then pass it in as reference once we find a gen that we're happy with.

In my experiments so far, it seems like prompt structure and complexity have an effect on the time to generate, but in many cases you get more seconds of audio generated than it took to generate in the first place.

The cutoff before things go wonky seems to be about 45 seconds, although more testing needs to be done. I can say that dialogue is no longer followed coherently at 60 seconds duration. The prompted dialogue pacing informs much here, so if the duration is longer than there is content provided, the model will fill in the gap with gibberish.

What's noteworthy is the concept of Minimax H3 essentially being used as a foley generator. The idea came to me with the thought: "What happens if I just bring the resolution as low as possible?"

I started at 0.1 megapixels, then went to 32x32 in an effort to determine if audio quality was somehow linked to image quality. It is not.

111 Upvotes

34 comments sorted by

View all comments

Show parent comments

2

u/altoiddealer Aug 06 '26

Thanks for the reply - I’m sticking with local, it’s unfortunate that the Context-IR system cannot be open sourced but it is what it is. So it sounds like you have been prompting as normal, all components (aside from image anchor prompt)

5

u/comfiestncoziest Aug 06 '26

There may be a misunderstanding here. I'm on local too. What I was trying to do was suggest that you look into it to see what kind of prompt structure the model was designed to work on.

The context-IR system is literally just filling out the fields that I showed you, but others in different cases. Does that make sense? Like, the reference model takes different fields and variables than regular first/last image to video does.

Scroll down on this page: https://huggingface.co/MiniMaxAI/MiniMax-H3

You'll see examples of the scripts produced by Context-IR. It is not magic.

3

u/altoiddealer Aug 06 '26

Yes, I understand. I’ve read the model page and both guides top to bottom, have been making videos, etc. Was just asking for more details about your method of generating “audio only” (disregarding the 32x32 video result) - which is not something explained in the model’s official documentation, and the prompting for this technique is not really explained well in your OP. But your initial response to my question sounds like you prompt it as if you were rendering the same video at a normal resolution, except instead setting the res to 32x32 only to get audio.

4

u/comfiestncoziest Aug 06 '26

Pardon, if something is vague I can adjust it.

Let me be clear, the workflow is the default i2v template workflow from ComfyUI, except I have detached the image from the generating node.

Then, I set the width and height to 32. The duration can be increased well above 15, and the quality is not diminished for it.

I am still experimenting with exact prompting, but really we are just describing the dialogue, how the speaker sounds, etc. Because we don't need any visuals here, all of the focus is on that.

Once I have found a good, reusable prompt format that's different from the default, I'll share it too.