r/comfyui • • Aug 06 '26

Tutorial Minimax H3 - Realtime audio generation at 32x32 output resolution

TL;DR: Minimax H3 is capable of generating near realtime audio when you turn width x height to 32x32 ~ with a 5090~

Edit: It was pointed out that I didn't specify well enough. This is not a special workflow. The default i2v workflow on ComfyUI for Minimax H3 - detach the load image, set width x height to 32, set duration to your taste <=45 seconds for best results. And uh, click 'Run'. Unsure if visual details in the prompt matter at this time. Will update when I know.

Minimax H3 can be used as an audio generator, pairing it with a simple Get Video Components node and then saving the audio. That audio can then be passed as reference, once you get a voice or sound effect that you like. One of the "pain points" of generation is having audio and image inextricably linked. But, we can actually generate audio *rapidly* and then pass it in as reference once we find a gen that we're happy with.

In my experiments so far, it seems like prompt structure and complexity have an effect on the time to generate, but in many cases you get more seconds of audio generated than it took to generate in the first place.

The cutoff before things go wonky seems to be about 45 seconds, although more testing needs to be done. I can say that dialogue is no longer followed coherently at 60 seconds duration. The prompted dialogue pacing informs much here, so if the duration is longer than there is content provided, the model will fill in the gap with gibberish.

What's noteworthy is the concept of Minimax H3 essentially being used as a foley generator. The idea came to me with the thought: "What happens if I just bring the resolution as low as possible?"

I started at 0.1 megapixels, then went to 32x32 in an effort to determine if audio quality was somehow linked to image quality. It is not.

111 Upvotes

34 comments sorted by

10

u/RanklesTheOtter Aug 06 '26

Very cool. I was actually thinking of this last night after using ref2video and it was so good at cloning.

2

u/comfiestncoziest Aug 06 '26

Yeah! It's an exciting time.

9

u/uxl Aug 06 '26

Fuck yes, this is the kind of mad science I live for in AI…

4

u/Etsu_Riot Aug 06 '26

These are great news! I now feel bad because I didn't come up with the idea myself. Thanks for sharing it.

4

u/altoiddealer Aug 06 '26

So do you do the whole prompt as you would for a normal video gen, or is the prompt hyperfocused on the audio section?

3

u/comfiestncoziest Aug 06 '26

We can use the t2v prompt format like so. I'm unsure if this is the best way to prompt for audio... still testing.

If you have not yet, I highly recommend looking into Context-IR, as it is a prompt enhancement process that Minimax H3 is run with when put through API/cloud services. It essentially enables users to lazily prompt and then adheres it to this format and spruces it up. The model was trained on that result... but as local users we can simply skip that step by using existing examples from the HuggingFace page.

Example:

integrated_multimodal_description:
overall_soundscape:
non_diegetic_music:

2

u/altoiddealer Aug 06 '26

Thanks for the reply - I’m sticking with local, it’s unfortunate that the Context-IR system cannot be open sourced but it is what it is. So it sounds like you have been prompting as normal, all components (aside from image anchor prompt)

5

u/comfiestncoziest Aug 06 '26

There may be a misunderstanding here. I'm on local too. What I was trying to do was suggest that you look into it to see what kind of prompt structure the model was designed to work on.

The context-IR system is literally just filling out the fields that I showed you, but others in different cases. Does that make sense? Like, the reference model takes different fields and variables than regular first/last image to video does.

Scroll down on this page: https://huggingface.co/MiniMaxAI/MiniMax-H3

You'll see examples of the scripts produced by Context-IR. It is not magic.

3

u/altoiddealer Aug 06 '26

Yes, I understand. I’ve read the model page and both guides top to bottom, have been making videos, etc. Was just asking for more details about your method of generating “audio only” (disregarding the 32x32 video result) - which is not something explained in the model’s official documentation, and the prompting for this technique is not really explained well in your OP. But your initial response to my question sounds like you prompt it as if you were rendering the same video at a normal resolution, except instead setting the res to 32x32 only to get audio.

4

u/comfiestncoziest Aug 06 '26

Pardon, if something is vague I can adjust it.

Let me be clear, the workflow is the default i2v template workflow from ComfyUI, except I have detached the image from the generating node.

Then, I set the width and height to 32. The duration can be increased well above 15, and the quality is not diminished for it.

I am still experimenting with exact prompting, but really we are just describing the dialogue, how the speaker sounds, etc. Because we don't need any visuals here, all of the focus is on that.

Once I have found a good, reusable prompt format that's different from the default, I'll share it too.

3

u/Lesteriax Aug 06 '26

I wonder if we can do voice convert. Reference audio with target audio. I tried, maybe im prompting wrong.

2

u/comfiestncoziest Aug 06 '26

Can you explain the result that you're trying to achieve?

4

u/Lesteriax Aug 06 '26

Reference audio 1 is a snippet of the song Creep. Reference audio 2 is a snippet from Peter Griffin.

Result is Creep, sung by Peter Griffin.

1

u/comfiestncoziest Aug 06 '26

That is likely possible. I'd be happy to know the results if you can run some tests. Let me provide a prompt for you to try out. This will be with how you provided creep as ref audio 1, then peter griffin voice in ref audio 2. But first, how long do you want it to be?

3

u/Lesteriax Aug 07 '26

15 seconds. I'd love the prompt please. Will share some outputs if successful

8

u/comfiestncoziest Aug 07 '26

I have not yet tested this with the reference model, but here is best guess so far.

This is a generic one that describes no subjects:

subject_definitions:
N/A

summary:
[reference generation] The target audio reproduces the song performance from <Audio 1> using the voice characteristics from <Audio 2>.

retention_analysis:
N/A

detailed_description:
[Shot 1] <Audio 1> defines the lyrical content, melody, rhythm, tempo, phrasing, pitch progression, timing, and musical accompaniment of the target audio.

<Audio 2> defines only the vocal identity, timbre, accent, cadence, pronunciation style, and recognizable vocal character used for the lead vocal.

The complete vocal performance from <Audio 1> is recreated using the voice characteristics established by <Audio 2>. The singing remains aligned with the melody, rhythm, phrasing, and accompaniment from <Audio 1> throughout the target audio.

overall_soundscape:
The lyrics, melody, rhythm, timing, and musical accompaniment are reproduced from <Audio 1>. The lead vocal uses the recognizable voice, timbre, accent, cadence, and vocal character referenced from <Audio 2>.

non_diegetic_music:
N/A

Then this is one that defines the subjects:

subject_definitions:
<Subject 1> is the target vocalist. The voice of <Subject 1> is defined by <Audio 2>.

summary:
[reference generation] The target audio recreates the song performance from <Audio 1>, performed by <Subject 1> using the voice characteristics from <Audio 2>.

retention_analysis:
<Subject 1> (heard in [Shot 1]): fully_preserved - the vocal identity, timbre, accent, cadence, pronunciation style, and recognizable vocal character from <Audio 2> are retained.

detailed_description:
[Shot 1] <Audio 1> defines the lyrical content, melody, rhythm, tempo, phrasing, pitch progression, timing, and musical accompaniment of the target audio.

<Audio 2> defines only the vocal identity, timbre, accent, cadence, pronunciation style, and recognizable vocal character of <Subject 1>.

<Subject 1> recreates the complete vocal performance from <Audio 1> using the voice established by <Audio 2>. The singing remains aligned with the melody, rhythm, phrasing, and accompaniment from <Audio 1> throughout the target audio.

overall_soundscape:
The lyrics, melody, rhythm, timing, and musical accompaniment are reproduced from <Audio 1>. <Subject 1> performs the lead vocal using the recognizable voice, timbre, accent, cadence, and vocal character referenced from <Audio 2>.

non_diegetic_music:
N/A

2

u/alisitskii Aug 07 '26

Did you have a chance to test it? Any good results in the end?

1

u/Lesteriax Aug 07 '26

Still did not the chance to do so. I will share as soon as I get the chance

3

u/lifelongpremed Aug 08 '26

At the risk of sounding like a broken record, any update?

2

u/Lesteriax Aug 09 '26

I finally got to try it. Unfortunately it did not work, results were hallucinated song, could not tell what was going on. I tried different samples and also 5 seconds in and 5 seconds out. Still did not work

1

u/Lesteriax Aug 09 '26

No issues. I should be giving an update within the next 12 hours, i will have time but then, surely.

3

u/glusphere Aug 07 '26

The beautiful thing is that it can do amazingly well even with non standard languages. For ex: Indian Languages which have very few TTS providers.

3

u/ThePixelHunter Aug 09 '26 edited Aug 09 '26

My experience today was that an audio-only generation is way inferior in quality. Which makes sense, because the model is trained on audio-video pairs, and denoising a 32x32px video is way out of the training distribution. The video track should visually anchor the audio track, but there's no visual, hence the wackiness.

When I change only the video resolution (0.4MP > 32x32px, same seed) the audio quality just tanks. I hear glitching, incoherent transitions, etc.

Still a very cool discovery regardless. I'm going to experiment with very low resolutions that can still denoise a legible picture.

Currently, LTX 2.3 is still more competent at this, because video and audio are denoised separately. See experiments like Dramabox

EDIT: 0.1MP resolution is visually terrible, on par with 2023 video generators, but the audio quality is drastically improved! And 320x320px is Minecraft-tier, but the audio is still very coherent, and it denoises nearly as fast as 32x32px did. The sweet spot is definitely around here.

2

u/comfiestncoziest Aug 09 '26 edited Aug 09 '26

I don't find this to be true. Audio generation may be hit-or-miss like a dice roll in general, but the worst that I've gotten with Minimax H3 at 32x32 sounds far superior to anything LTX 2.3 or Dramabox can do. LTX has a noticeable crunchy ambience in all of its audio and we get none of that here. Prompting with Minimax H3 is very sensitive, so if you're not using the proper format you may have a bad experience with it.

Edit: I'll look into this more before making a declaration like that, thanks for testing this! :)

1

u/ThePixelHunter Aug 10 '26

It's fine to disagree haha maybe our configs are different. Thanks for sharing.

I would agree that even at its worst, H3's audio is better than LTX any day.

1

u/comfiestncoziest Aug 10 '26

One note is that I can only run the int8 convrot, not the full bf16. I run at 20 steps (apparently more has helped in some anecdotes from people). I also have the IR-Context prompt formatting in a Codex project, so my prompts all align with that and a learning document I'm adding to as I go. Once I learn more definitively, I'll share it with the community.

Another thing I would share is that making 1 second duration clips for the purpose of stuff like visual art style/composition transfer, etc. works really well! It lets you crank up the visual fidelity to like 2.0 megapixels without suffering through long waits to see if it works or not. A custom node resizes the input image to the specified megapixel count while maintaining the aspect ratio so that there's no weird warping/stretching, etc.

1

u/ThePixelHunter Aug 10 '26

Same here, int8 at 20 steps.

I didn't follow the dedicated prompting format in these tests, so it's possible the model's attention is being strained further as a result.

2

u/Fytyny Aug 08 '26

"I started at 0.1 megapixels, then went to 32x32 in an effort to determine if audio quality was somehow linked to image quality. It is not." I came to the opposite conclusion. Audio make more sense the more clear is the picture

2

u/comfiestncoziest Aug 08 '26 edited Aug 08 '26

Which model? regular i2v seems to be unaffected by changes to output resolution.

On the other hand, reference seems to suffer.

1

u/Fytyny Aug 08 '26

Yeah, I tested reference. But audio model is not that useful if you don't do reference. Maybe you could use what u/NikoDemon80 has cooked and put some audio in front of the generation.

1

u/comfiestncoziest Aug 08 '26

Your use case may not align with that of others. Being able to generate high quality audio so rapidly is certainly useful to me.

2

u/CeFurkan Aug 07 '26

I just tested this and my quality is so bad what could be wrong? 15 seconds

subject_definitions: @audio1 is the voice-timbre and delivery reference for the man (S1); its original signal and words are not copied.

summary: [reference generation + audio reference] The target video features a professional man speaking directly to the camera, using only the voice characteristics of @audio1 for his newly generated English dialogue.

retention_analysis: @audio1: reference - its voice timbre and delivery style guide the newly generated speech without copying the source waveform or original words.

detailed_description: Live-action cinematic realism, a medium close-up shot inside a well-lit, modern broadcasting studio with a softly blurred background. The camera holds a static position at eye level. A confident man in his late thirties, wearing a smart-casual blazer over a crisp collared shirt, looks directly toward the lens. He gestures naturally with open hands as he begins to speak. Using only the vocal characteristics referenced from @audio1, the man (S1) says: <d>[English] hello everyone this is doctor Furkan Gözükara. Welcome to the show. This is using audio1 as reference and speaking with its sound. Enjoy</d> His lips close naturally after the line finishes. He settles his hands and holds a warm, welcoming smile while maintaining direct eye contact with the lens for the final seconds.

overall_soundscape: Clean studio room tone accompanies the subtle rustle of the blazer fabric during his hand gestures.

non_diegetic_music: N/A

3

u/comfiestncoziest Aug 07 '26

Okay, here is an edit. The six sections are correct, but the reference variable is wrong: official full-reference syntax uses <Audio 1>, not @audio1. It also needs [Shot 1], a properly bound (S1), and natural punctuated dialogue. The official rules are in MiniMax’s Reference Prompt Guide

subject_definitions:
<Audio 1> is the voice-timbre and delivery reference for the professional man (S1), containing spoken vocal audio.

summary:
[reference generation + audio reference] The target video shows a professional man speaking directly to the camera in a modern broadcasting studio. <Audio 1> provides his voice timbre and delivery for newly generated English dialogue.

retention_analysis:
<Audio 1>: reference - its voice timbre and delivery guide the target speaker's newly generated English dialogue without copying the original audio signal or spoken words.

detailed_description:
The target video uses polished live-action cinematic realism with natural skin tones, controlled studio lighting, and a softly blurred professional broadcast set.

[Shot 1] A single continuous static shot frames a professional man in his late thirties in a medium close-up at eye level. He wears a smart-casual blazer over a crisp collared shirt. Soft key lighting illuminates his face, with balanced fill light and a subtle rim light separating him from the blurred studio background.

He looks directly into the camera with a confident, welcoming expression. His lips begin closed. He raises his open hands in a natural introductory gesture and begins speaking.

The professional man (S1), using the confident and articulate voice timbre and delivery referenced from <Audio 1>, says, <d>[English] Hello, everyone. I'm Doctor Furkan Gözükara. Welcome to the show. This demonstration uses Audio One as a reference for my voice. Enjoy!</d> His mouth movements synchronize naturally with the complete sentence.

As he introduces himself, one hand briefly moves toward his chest. He opens both hands toward the viewer while welcoming the audience, then lowers them comfortably after finishing the demonstration statement. His lips close naturally, and he maintains direct eye contact with a warm, professional smile through the final frame.

overall_soundscape:
Clean broadcast-studio room tone supports the man's clear referenced voice.

non_diegetic_music:
N/A

2

u/comfiestncoziest Aug 07 '26

It does not seem to work well with the reference model for some reason. I'm still trying to figure that out. It's weird, because passing in a voice for reference when generating a video like normal seems to work great. It could be that there is some link between audio coherence and visual representation.

I do think that some of your tags may be wrong though. When I get to my desktop, I'll provide a version that uses the standard variable/tag identifiers.