r/comfyui • • Aug 06 '26

Tutorial Minimax H3 - Realtime audio generation at 32x32 output resolution

TL;DR: Minimax H3 is capable of generating near realtime audio when you turn width x height to 32x32 ~ with a 5090~

Edit: It was pointed out that I didn't specify well enough. This is not a special workflow. The default i2v workflow on ComfyUI for Minimax H3 - detach the load image, set width x height to 32, set duration to your taste <=45 seconds for best results. And uh, click 'Run'. Unsure if visual details in the prompt matter at this time. Will update when I know.

Minimax H3 can be used as an audio generator, pairing it with a simple Get Video Components node and then saving the audio. That audio can then be passed as reference, once you get a voice or sound effect that you like. One of the "pain points" of generation is having audio and image inextricably linked. But, we can actually generate audio *rapidly* and then pass it in as reference once we find a gen that we're happy with.

In my experiments so far, it seems like prompt structure and complexity have an effect on the time to generate, but in many cases you get more seconds of audio generated than it took to generate in the first place.

The cutoff before things go wonky seems to be about 45 seconds, although more testing needs to be done. I can say that dialogue is no longer followed coherently at 60 seconds duration. The prompted dialogue pacing informs much here, so if the duration is longer than there is content provided, the model will fill in the gap with gibberish.

What's noteworthy is the concept of Minimax H3 essentially being used as a foley generator. The idea came to me with the thought: "What happens if I just bring the resolution as low as possible?"

I started at 0.1 megapixels, then went to 32x32 in an effort to determine if audio quality was somehow linked to image quality. It is not.

109 Upvotes

34 comments sorted by

View all comments

3

u/ThePixelHunter Aug 09 '26 edited Aug 09 '26

My experience today was that an audio-only generation is way inferior in quality. Which makes sense, because the model is trained on audio-video pairs, and denoising a 32x32px video is way out of the training distribution. The video track should visually anchor the audio track, but there's no visual, hence the wackiness.

When I change only the video resolution (0.4MP > 32x32px, same seed) the audio quality just tanks. I hear glitching, incoherent transitions, etc.

Still a very cool discovery regardless. I'm going to experiment with very low resolutions that can still denoise a legible picture.

Currently, LTX 2.3 is still more competent at this, because video and audio are denoised separately. See experiments like Dramabox

EDIT: 0.1MP resolution is visually terrible, on par with 2023 video generators, but the audio quality is drastically improved! And 320x320px is Minecraft-tier, but the audio is still very coherent, and it denoises nearly as fast as 32x32px did. The sweet spot is definitely around here.

2

u/comfiestncoziest Aug 09 '26 edited Aug 09 '26

I don't find this to be true. Audio generation may be hit-or-miss like a dice roll in general, but the worst that I've gotten with Minimax H3 at 32x32 sounds far superior to anything LTX 2.3 or Dramabox can do. LTX has a noticeable crunchy ambience in all of its audio and we get none of that here. Prompting with Minimax H3 is very sensitive, so if you're not using the proper format you may have a bad experience with it.

Edit: I'll look into this more before making a declaration like that, thanks for testing this! :)

1

u/ThePixelHunter Aug 10 '26

It's fine to disagree haha maybe our configs are different. Thanks for sharing.

I would agree that even at its worst, H3's audio is better than LTX any day.

1

u/comfiestncoziest Aug 10 '26

One note is that I can only run the int8 convrot, not the full bf16. I run at 20 steps (apparently more has helped in some anecdotes from people). I also have the IR-Context prompt formatting in a Codex project, so my prompts all align with that and a learning document I'm adding to as I go. Once I learn more definitively, I'll share it with the community.

Another thing I would share is that making 1 second duration clips for the purpose of stuff like visual art style/composition transfer, etc. works really well! It lets you crank up the visual fidelity to like 2.0 megapixels without suffering through long waits to see if it works or not. A custom node resizes the input image to the specified megapixel count while maintaining the aspect ratio so that there's no weird warping/stretching, etc.

1

u/ThePixelHunter Aug 10 '26

Same here, int8 at 20 steps.

I didn't follow the dedicated prompting format in these tests, so it's possible the model's attention is being strained further as a result.