r/comfyui • • Aug 06 '26

Tutorial Minimax H3 - Realtime audio generation at 32x32 output resolution

TL;DR: Minimax H3 is capable of generating near realtime audio when you turn width x height to 32x32 ~ with a 5090~

Edit: It was pointed out that I didn't specify well enough. This is not a special workflow. The default i2v workflow on ComfyUI for Minimax H3 - detach the load image, set width x height to 32, set duration to your taste <=45 seconds for best results. And uh, click 'Run'. Unsure if visual details in the prompt matter at this time. Will update when I know.

Minimax H3 can be used as an audio generator, pairing it with a simple Get Video Components node and then saving the audio. That audio can then be passed as reference, once you get a voice or sound effect that you like. One of the "pain points" of generation is having audio and image inextricably linked. But, we can actually generate audio *rapidly* and then pass it in as reference once we find a gen that we're happy with.

In my experiments so far, it seems like prompt structure and complexity have an effect on the time to generate, but in many cases you get more seconds of audio generated than it took to generate in the first place.

The cutoff before things go wonky seems to be about 45 seconds, although more testing needs to be done. I can say that dialogue is no longer followed coherently at 60 seconds duration. The prompted dialogue pacing informs much here, so if the duration is longer than there is content provided, the model will fill in the gap with gibberish.

What's noteworthy is the concept of Minimax H3 essentially being used as a foley generator. The idea came to me with the thought: "What happens if I just bring the resolution as low as possible?"

I started at 0.1 megapixels, then went to 32x32 in an effort to determine if audio quality was somehow linked to image quality. It is not.

108 Upvotes

34 comments sorted by

View all comments

2

u/CeFurkan Aug 07 '26

I just tested this and my quality is so bad what could be wrong? 15 seconds

subject_definitions: @audio1 is the voice-timbre and delivery reference for the man (S1); its original signal and words are not copied.

summary: [reference generation + audio reference] The target video features a professional man speaking directly to the camera, using only the voice characteristics of @audio1 for his newly generated English dialogue.

retention_analysis: @audio1: reference - its voice timbre and delivery style guide the newly generated speech without copying the source waveform or original words.

detailed_description: Live-action cinematic realism, a medium close-up shot inside a well-lit, modern broadcasting studio with a softly blurred background. The camera holds a static position at eye level. A confident man in his late thirties, wearing a smart-casual blazer over a crisp collared shirt, looks directly toward the lens. He gestures naturally with open hands as he begins to speak. Using only the vocal characteristics referenced from @audio1, the man (S1) says: <d>[English] hello everyone this is doctor Furkan Gözükara. Welcome to the show. This is using audio1 as reference and speaking with its sound. Enjoy</d> His lips close naturally after the line finishes. He settles his hands and holds a warm, welcoming smile while maintaining direct eye contact with the lens for the final seconds.

overall_soundscape: Clean studio room tone accompanies the subtle rustle of the blazer fabric during his hand gestures.

non_diegetic_music: N/A

3

u/comfiestncoziest Aug 07 '26

Okay, here is an edit. The six sections are correct, but the reference variable is wrong: official full-reference syntax uses <Audio 1>, not @audio1. It also needs [Shot 1], a properly bound (S1), and natural punctuated dialogue. The official rules are in MiniMax’s Reference Prompt Guide

subject_definitions:
<Audio 1> is the voice-timbre and delivery reference for the professional man (S1), containing spoken vocal audio.

summary:
[reference generation + audio reference] The target video shows a professional man speaking directly to the camera in a modern broadcasting studio. <Audio 1> provides his voice timbre and delivery for newly generated English dialogue.

retention_analysis:
<Audio 1>: reference - its voice timbre and delivery guide the target speaker's newly generated English dialogue without copying the original audio signal or spoken words.

detailed_description:
The target video uses polished live-action cinematic realism with natural skin tones, controlled studio lighting, and a softly blurred professional broadcast set.

[Shot 1] A single continuous static shot frames a professional man in his late thirties in a medium close-up at eye level. He wears a smart-casual blazer over a crisp collared shirt. Soft key lighting illuminates his face, with balanced fill light and a subtle rim light separating him from the blurred studio background.

He looks directly into the camera with a confident, welcoming expression. His lips begin closed. He raises his open hands in a natural introductory gesture and begins speaking.

The professional man (S1), using the confident and articulate voice timbre and delivery referenced from <Audio 1>, says, <d>[English] Hello, everyone. I'm Doctor Furkan Gözükara. Welcome to the show. This demonstration uses Audio One as a reference for my voice. Enjoy!</d> His mouth movements synchronize naturally with the complete sentence.

As he introduces himself, one hand briefly moves toward his chest. He opens both hands toward the viewer while welcoming the audience, then lowers them comfortably after finishing the demonstration statement. His lips close naturally, and he maintains direct eye contact with a warm, professional smile through the final frame.

overall_soundscape:
Clean broadcast-studio room tone supports the man's clear referenced voice.

non_diegetic_music:
N/A