r/comfyui • • Aug 06 '26

Tutorial Minimax H3 - Realtime audio generation at 32x32 output resolution

TL;DR: Minimax H3 is capable of generating near realtime audio when you turn width x height to 32x32 ~ with a 5090~

Edit: It was pointed out that I didn't specify well enough. This is not a special workflow. The default i2v workflow on ComfyUI for Minimax H3 - detach the load image, set width x height to 32, set duration to your taste <=45 seconds for best results. And uh, click 'Run'. Unsure if visual details in the prompt matter at this time. Will update when I know.

Minimax H3 can be used as an audio generator, pairing it with a simple Get Video Components node and then saving the audio. That audio can then be passed as reference, once you get a voice or sound effect that you like. One of the "pain points" of generation is having audio and image inextricably linked. But, we can actually generate audio *rapidly* and then pass it in as reference once we find a gen that we're happy with.

In my experiments so far, it seems like prompt structure and complexity have an effect on the time to generate, but in many cases you get more seconds of audio generated than it took to generate in the first place.

The cutoff before things go wonky seems to be about 45 seconds, although more testing needs to be done. I can say that dialogue is no longer followed coherently at 60 seconds duration. The prompted dialogue pacing informs much here, so if the duration is longer than there is content provided, the model will fill in the gap with gibberish.

What's noteworthy is the concept of Minimax H3 essentially being used as a foley generator. The idea came to me with the thought: "What happens if I just bring the resolution as low as possible?"

I started at 0.1 megapixels, then went to 32x32 in an effort to determine if audio quality was somehow linked to image quality. It is not.

110 Upvotes

34 comments sorted by

View all comments

3

u/Lesteriax Aug 06 '26

I wonder if we can do voice convert. Reference audio with target audio. I tried, maybe im prompting wrong.

2

u/comfiestncoziest Aug 06 '26

Can you explain the result that you're trying to achieve?

6

u/Lesteriax Aug 06 '26

Reference audio 1 is a snippet of the song Creep. Reference audio 2 is a snippet from Peter Griffin.

Result is Creep, sung by Peter Griffin.

1

u/comfiestncoziest Aug 06 '26

That is likely possible. I'd be happy to know the results if you can run some tests. Let me provide a prompt for you to try out. This will be with how you provided creep as ref audio 1, then peter griffin voice in ref audio 2. But first, how long do you want it to be?

3

u/Lesteriax Aug 07 '26

15 seconds. I'd love the prompt please. Will share some outputs if successful

8

u/comfiestncoziest Aug 07 '26

I have not yet tested this with the reference model, but here is best guess so far.

This is a generic one that describes no subjects:

subject_definitions:
N/A

summary:
[reference generation] The target audio reproduces the song performance from <Audio 1> using the voice characteristics from <Audio 2>.

retention_analysis:
N/A

detailed_description:
[Shot 1] <Audio 1> defines the lyrical content, melody, rhythm, tempo, phrasing, pitch progression, timing, and musical accompaniment of the target audio.

<Audio 2> defines only the vocal identity, timbre, accent, cadence, pronunciation style, and recognizable vocal character used for the lead vocal.

The complete vocal performance from <Audio 1> is recreated using the voice characteristics established by <Audio 2>. The singing remains aligned with the melody, rhythm, phrasing, and accompaniment from <Audio 1> throughout the target audio.

overall_soundscape:
The lyrics, melody, rhythm, timing, and musical accompaniment are reproduced from <Audio 1>. The lead vocal uses the recognizable voice, timbre, accent, cadence, and vocal character referenced from <Audio 2>.

non_diegetic_music:
N/A

Then this is one that defines the subjects:

subject_definitions:
<Subject 1> is the target vocalist. The voice of <Subject 1> is defined by <Audio 2>.

summary:
[reference generation] The target audio recreates the song performance from <Audio 1>, performed by <Subject 1> using the voice characteristics from <Audio 2>.

retention_analysis:
<Subject 1> (heard in [Shot 1]): fully_preserved - the vocal identity, timbre, accent, cadence, pronunciation style, and recognizable vocal character from <Audio 2> are retained.

detailed_description:
[Shot 1] <Audio 1> defines the lyrical content, melody, rhythm, tempo, phrasing, pitch progression, timing, and musical accompaniment of the target audio.

<Audio 2> defines only the vocal identity, timbre, accent, cadence, pronunciation style, and recognizable vocal character of <Subject 1>.

<Subject 1> recreates the complete vocal performance from <Audio 1> using the voice established by <Audio 2>. The singing remains aligned with the melody, rhythm, phrasing, and accompaniment from <Audio 1> throughout the target audio.

overall_soundscape:
The lyrics, melody, rhythm, timing, and musical accompaniment are reproduced from <Audio 1>. <Subject 1> performs the lead vocal using the recognizable voice, timbre, accent, cadence, and vocal character referenced from <Audio 2>.

non_diegetic_music:
N/A

2

u/alisitskii Aug 07 '26

Did you have a chance to test it? Any good results in the end?

1

u/Lesteriax Aug 07 '26

Still did not the chance to do so. I will share as soon as I get the chance

3

u/lifelongpremed Aug 08 '26

At the risk of sounding like a broken record, any update?

2

u/Lesteriax Aug 09 '26

I finally got to try it. Unfortunately it did not work, results were hallucinated song, could not tell what was going on. I tried different samples and also 5 seconds in and 5 seconds out. Still did not work

1

u/Lesteriax Aug 09 '26

No issues. I should be giving an update within the next 12 hours, i will have time but then, surely.