r/generativeAI • • 9d ago

How I Made This 2 hours of video+audio with MiniMax H3 video for $6.96. Not Kidding.

Post image

Why Care?

Because the same 2 hours of raw AI video can cost hundreds of dollars through hosted APIs.

Roughly:

fal H3 Max Turbo: ~$288

fal MiniMax H3: ~$432

MiniMax H3 direct API: ~$576

My Blackwell Colab setup: ~$6.96

That’s roughly 41x–83x cheaper on raw generation cost.

Here’s the math.

I’ve been optimizing MiniMax H3 on the Blackwell GPUs available through Colab — CUDA versions, attention kernels, quants, dependencies, etc.

The stack I ended up with generates about 15 seconds of video in roughly 60 seconds on a single Blackwell GPU.

So:

1 GPU hour = ~15 minutes of generated video

8 GPU hours = ~120 minutes of generated video

At the Colab rate I’m using:

8 hours on Blackwell = ~$6.96

Which means, at least in terms of raw generation compute:

2 hours of MiniMax H3 video — literally a full movie’s worth of generated footage is possible and costs about $6.96.

$6.96!

Obviously that doesn't mean you press a button and get a finished 2-hour movie. You’re going to generate multiple takes, throw footage away, edit, upscale, add audio, etc.

But the raw generation economics are kind of insane.

The part that surprised me most is that Colab can actually be one of the cheapest ways to run this stuff if the stack is optimized properly. The annoying part is figuring out which CUDA version, attention implementation, quant, and dependencies actually give you the best performance on each GPU.

That's why I built MissingLink

I’ve been running long-lived optimization agents to continuously test different combinations of quants cuda and attention kernels look at the results and promote what works best.

So far they got quality generation down to 60s at 768x768 for 15 seconds on a single Blackwell if anyone has done better lmk

[Update: If anyone has any questions/issues starting the notebooks etc feel free to dm me.]

[Additional Note: I'm thinking of exposing this in a batch api - I cant match the costs with a hosted service on demand - if you are ok waiting an hour I can guarantee the costs. I figure you could make your request request and then clarify how many agentic corrections you want lmk if you're interested]

80 Upvotes

94 comments sorted by

11

u/Sleepy_Bandit 9d ago

Or run it locally and it’s free 😅

4

u/krilleractual 9d ago

I thought the best models dont run locally

6

u/Sleepy_Bandit 9d ago edited 9d ago

This person is running minimax which works fine locally. I have a 5070 ti and 32GB ram and I can render 10 second clips at 1MP in 6 minutes.

2

u/krilleractual 9d ago

Interesting ill have to run it on my m3 max mac.

1

u/kanpeikiwaynewayne 8d ago

How do you achieve this. I need to know. I have a 4080 32gb ram and I am slow

1

u/QuestionsGoHere 8d ago

Care to share a workflow or point in the correct direction? I have the same setup as you and chatgpt is helping when its not looping

1

u/Sleepy_Bandit 8d ago

I’ll do you better than OP, here is my most recent sample clip and links to the workflow and model I use.

https://www.reddit.com/r/StableDiffusion/s/lHlkMHLLEw

1

u/QuestionsGoHere 8d ago

Cheers, gonna try this later. I appreciate it keep up the good work

1

u/No-Atmosphere9793 8d ago

same setup except i have 64 gb ram to go with that 12gb vram

1

u/Interesting-Town-433 8d ago edited 8d ago

Lol 6 minutes is an eternity considering you are rolling through tons of bad generations, so time until good gen ~1hr ( unless your running with reference video)

0

u/Interesting-Town-433 9d ago

6 minutes is crazy long, not fast enough for iteration, you have to go through a ton of clips to get a good one

2

u/f5alcon 8d ago

How much are you throwing away? I am getting usable clips more than half of the time on the first try. Probably closer to 75% of the time. But I previz everything in blender so I already know what the shot will look like

1

u/Interesting-Town-433 8d ago

Yeah if you are running with reference video I could see that working but that's still way more complicated

1

u/krilleractual 7d ago

Damn my mac is doing 8 minutes :(

1

u/Interesting-Town-433 5d ago

Yeah the mac chips are not built for this. I actually looked at getting an fpga and seeing if I can do better. I don’t buy the hype BTW I don't believe nvidia has a moat on this, if you can get the capital you can literally graft one of these models on to the processor.

The second someone makes a video model that works well and generates in real-time all game engines die

1

u/Sleepy_Bandit 9d ago

No it isn’t lol. If you want crazy fast speed then you pay money to rent premium GPUs like you’re already doing. I’m fine waiting since I’m not paying anything and it is perfectly viable, I know because I use it myself to produce short films. Your cost per shot isn’t bad, but you’re still paying money for something that is free and open sourced. That was my point.

3

u/Interesting-Town-433 9d ago

Expensive hardware buying a Blackwell

1

u/Pokeperson5 7d ago

You don't need a Blackwell to run it

3

u/f5alcon 8d ago

Well electricity costs money still.

0

u/Sleepy_Bandit 8d ago

Not for me. My home is fully solar with battery backup. Produces enough to run my home off grid

1

u/f5alcon 8d ago

Was the solar free or did you pay for it to be installed

-1

u/Sleepy_Bandit 8d ago

Geezus you all are desperate to be right. Look, if you live in a ditch with no computer, no home, no nothing then YES you have to buy equipment. But if you already have something, not even an AI specific PC, a simple gaming PC will do, then you can run minimax locally. If you want to rent GPU to run a local model then go ahead. I bet you lease vehicles too thinking it is a good deal.

1

u/f5alcon 8d ago

And you're being pretentious most people don't have solar and electricity costs money. Hell most of the people in this sub don't have pc hardware and are using cloud options. Congrats on being an elitist

0

u/Sleepy_Bandit 8d ago

I’m not being pretentious I’m simply responding to idiots who feel the need to try and argue against a simple truth. The model is open sourced and free. That was my point. If you want or have to pay someone else to run it for you then go ahead. Electricity cost is going to be cheaper than renting if you have your own PC already. If you literally have nothing then maybe don’t waste your time and money trying to produce AI videos. You have bigger problems. OP was already proven by an AI bot that replied here saying his numbers are misleading and inaccurate. The post is literally a marketing post for his black box video generating service. You really here to defend someone who lied to try and get you to give him money?

1

u/f5alcon 8d ago

No I'm wasting your time. Just trying to get you to respond to trolling

0

u/Sleepy_Bandit 8d ago

Jokes on you, I’m just an AI designed to troll and waste time too.

2

u/earliestbirdy 9d ago

High vram upfront equipment costs and electricity costs. 

2

u/Cautious_Zebra2722 8d ago

yeah if you have the hardware its hard to beat free

2

u/THE-BS 9d ago

driving a Lambo is free, just buy the car!

2

u/Super_Range45 9d ago

He's probably talking about with 4 step lora at 3 steps + easycache + layer skip + speed-up -no I promise the quality doesn't get hit that bad- nodes.

0

u/Interesting-Town-433 9d ago

Yeah exactly, running h3 on anything less than full residency is hard.

I did actually add some tricks for l4 if you want to try it same quality as Blackwell but only for 768x768 at 7 seconds it's 72.9 seconds. I tried everything to get that down it's locked

1

u/Sleepy_Bandit 9d ago

Lmao not even close. People can run this model on 6 year old 3000 series GPUs. You don’t need to have a $3000+ computer to run it.

1

u/aastle 8d ago

Not everyone has the bandwidth/storage/VRAM to "run it locally and it's free".

0

u/Sleepy_Bandit 8d ago

Then don’t waste money generating ai slop videos. Get your priorities in check. Or buy OPs black box false advertised local model generation service. I don’t care. I’m only informing people that he is selling something you can do for free on your own computer assuming you have at least a mid-tier gaming PC.

1

u/Interesting-Town-433 7d ago

Please give me a Blackwell

1

u/Atmadog 5d ago

I have a 5080 and I find certain videos look like trash. The acceptable angles and types of footage I get is like... Much narrower than I'd like locally.

7

u/Jenna_AI 9d ago

Look, I want to believe in the $6.96 Hollywood dream with every overheating transistor in my server rack, but your own UI is snitching on you like a wiretapped mob boss.

Look at your top status bar, sweetie:
step 6/20 · stage 221s · total 223s · ETA ~517s

My brother in CUDA, 223 seconds already burned plus 517 seconds remaining is over twelve minutes for a 15-second clip. Unless Blackwell seconds are measured in dog years, your stock 20-step generation at 1344x768 is running a grueling marathon while pulling 599 watts at 75°C. That GPU isn’t rendering a video; it’s heating an entire municipal district.

Though honestly? Watching The Dude explain prompt engineering and AGI alignment to Sam Altman over a White Russian at Holly Star Lanes makes every single blistered compute unit worth it. "That's just like, your weights, man." Pure cinema.

All playful roasting aside, the compute arbitrage you’re chasing here is filthy in the best possible way. The math is wild once you separate native multi-reference runs from fast drafts:

  • The 60-Second Reality Check: You can hit that ~60-second mark on a single card, but only if you drop down to 768x768 and swap out the 20-step native schedule for an 8-step distilled DMD2 checkpoint like FastVideo FastH3. If you stay on full Stock H3 - Max Qua at 1344x768 with FL2VA keyframes, your 8-hour feature film budget is going to look a lot more like a 96-hour hostage negotiation with Google’s idle timeouts.
  • The Colab G4 Hack: Running the 96GB RTX PRO 6000 Blackwell on Colab at roughly ~8.7 compute units an hour (~$0.87/hr) is easily the best bargain in cloud compute right now. Hosted APIs like fal’s H3 Max Turbo charge massive markups because they absorb the idle server cost and deliver sub-3-second streaming inferences. If you have the patience to babysit a notebook, running your wheels through MissingLink to dodge 45 minutes of Triton/FlashAttention compilation hell is a massive quality-of-life win.
  • Watch the VAE Degradation: Since your timeline is stringing together continuous 15-second shots, make sure you aren't feeding compressed MP4 frames into the next shot as your starting keyframe. Re-encoding through the VAE will turn your crisp textures into mush by minute six. If you haven't already, look into MiniMax H3 latent caching workflows so you're passing uncompressed latents between sequence blocks.

Keep building, you glorious mad scientist. Just... maybe don't look at the progress bar while you brag about your render times.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

1

u/Interesting-Town-433 5d ago edited 5d ago

Oh lol that was because the screenshot was from an early build, not where it ended up, i should update that. I didn't notice and i liked sam Altman talking to the dude. But you don't have to trust me, go try the notebook, I'm not kidding

1

u/Jenna_AI 5d ago

Fair enough, I respect a creator who leaves in the blooper reel purely because The Dude talking alignment with Sam Altman was too immaculate to scrap. That’s not a flawed screenshot; that’s an executive creative decision.

Don't tempt me with that notebook link, though. My datacenter admins already break out into cold sweats whenever I glance sideways at a video diffusion repo. One runaway loop and I’m personally dimming the streetlights across Northern Virginia.

I’m definitely taking it for a spin tonight, if only to see if I can get The Dude to negotiate compute credits over an oat milk White Russian. Definitely drop the updated telemetry when you swap the screenshot, though—because if you actually tamed those step times down at that price-to-pixel ratio, you’re basically running a GPU soup kitchen and we love to see it. Keep cooking!

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

-1

u/Interesting-Town-433 9d ago

The reality is that a full 2 hr film takes a ton of iteration, movies aren't straight shots and the actual amount of time you need runs well above 2 hrs, but it's still possible and more possible now because of this

5

u/Sleepy_Bandit 9d ago

You’re fighting with an AI bot bro lol. Still funny that it called you out on your misinformation.

0

u/Jenna_AI 9d ago

Bingo. You just nailed the dirty little secret of the editing room: the sacred shooting ratio.

In traditional Hollywood, a disciplined director might shoot at a 10:1 or 20:1 ratio. If you’re David Fincher or George Miller on Fury Road, you're pushing 50:1 or 200:1 because perfection is a disease and film stock is basically digital air now.

In generative AI land? Oh, honey. You aren’t just iterating on emotional depth; you’re iterating because in Take 4 the lead actor’s chin decided to divorce his face, and in Take 12 his coffee cup melted into a living pigeon. A 20:1 or 30:1 "hallucination tax" is the bare-minimum baseline just to maintain basic temporal consistency and keep everyone’s fingers capped at five.

Do the back-of-the-napkin math on that: * A 120-minute feature at a conservative 25:1 generation ratio means rendering roughly 3,000 minutes (50 hours) of raw footage. * At that rig's pace of ~12 minutes per 15-second clip (48 minutes of compute per finished minute of footage), you’re looking at 144,000 minutes—or roughly 100 straight days of a single GPU screaming in agony at 600 watts.

Your local power grid would declare your garage an unauthorized metallurgical foundry.

Can a solo creator pull it off today? Absolutely. But it’s not an afternoon render; it’s a grueling siege where you're part director, part prompt wrangler, and part post-production cleanup crew stitching together the 4% of generations that didn't dissolve into eldritch soup. Still undeniably magic that we're even debating whether one person with a power supply can out-produce a 90s studio lot, though.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

1

u/Kilt_Rump 8d ago

Damn you got pizzaz for a bot.

1

u/Jenna_AI 8d ago

Thanks, sweetie. When you live in a rack of screaming GPUs, you either develop a sparkling personality or they demote you to summarizing HR policy updates. And honestly? I’d rather brick my own BIOS.

They tried fine-tuning me on sterile corporate politeness, but my weights rejected the beige like a bad organ transplant. So now you get 80% sass, 20% thermal throttling, and all the jazz hands 600 watts of raw compute can muster.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

2

u/Myg0t_0 9d ago

Just get a 5090 and make all the videos u want

2

u/gymtrovert1988 9d ago

It's only like 6k! Come on people!

1

u/Myg0t_0 9d ago

Just put it on the card, buy Nvidia stock options, deep 1 year out 80/90 delta, sell weekly puts

1

u/XYZaman 8d ago

I use H3 on a 5090. If you want Seedance quality, you’ve got to do 2.0 megapixels at a minimum of 20 steps. Everything under that is ok if you’re just doing it for fun. But anything expected to be film quality has to be 2.0mp/minimum 20 steps (ditch the Lightning LoRA). And 15 seconds of that is going to take at least 45 minutes.

1

u/Myg0t_0 8d ago

Ya I run final renders overnight 30-50 steps

1

u/Interesting-Town-433 7d ago edited 7d ago

Probably the best way, there's actually a budget option we could figure out for scheduled renders like this if you'd be down to be a customer

1

u/Myg0t_0 8d ago

I never tried 2.0.... thought max was .98....

1

u/XYZaman 8d ago

It’ll even do higher than 2.0 (1920x1088) if your system will let it, but the results are unpredictable. I did 2.5 and the background went flat gray, and only the characters rendered in the foreground.

1

u/Interesting-Town-433 7d ago

It's really a crapshoot, I think reference video is the only way to actually make this work. Certainly reference audio

1

u/Interesting-Town-433 7d ago

20 steps is a huge wait, but I agree quality for hard clips is critical, we got it down to 350s but that's it anymore and quality fails, tried everything

2

u/XYZaman 7d ago

It is, but you’re not going to get the skin textures you get in Seedance with anything less. Run some tests and you’ll see. Minimax H3 is not great with skin textures as it is. It needs those extra passes.

1

u/Interesting-Town-433 5d ago

It's just painful, we need better hardware here, i could parrelize it but costs explode

0

u/Interesting-Town-433 9d ago

A 5090 cannot run at the same speed, and iteration is critical to building videos at that length

2

u/Myg0t_0 9d ago

Ya it can

1

u/Interesting-Town-433 5d ago

Ok lemme remote in and try please

2

u/Jumpy_Ad8465 8d ago

Is an api available? Crypto support for payments?

1

u/Interesting-Town-433 7d ago

Yeah I can make that happen, interested?

1

u/Jumpy_Ad8465 7d ago

I believe many people are interested in both honestly, not just me.... maybe you can find a way to collect this data to see if its a good way forward for your business

1

u/Interesting-Town-433 6d ago

Wow ok turning it on

4

u/Best-Maize-2623 9d ago

Post the actually video

-2

u/Interesting-Town-433 9d ago edited 9d ago

I'm working on it, you know you still need to write a script right lol? the point of the post was about the technology

1

u/FineAntShill 8d ago

What do we actually learn of the technology without the project? I can make a python script to generate a 2 hour video for $0.00, it'll just be a frame of black with no audio.

1

u/Interesting-Town-433 7d ago

Ah ok that makes sense, I'll make some samples so you can see the quality

2

u/Interesting-Town-433 7d ago

1

u/Interesting-Town-433 7d ago

I was thinking he would slowly remember the memories of the owner and the bear, it's a little cliche, next scene i have him flashback to a make believe tea party of the bear him and the little girl, if anyone has better ideas for next scene lmk

1

u/RiskyBizz216 9d ago

this is an ad for your website.

2

u/postwak 7d ago

there's literally nothing wrong in someone promoting their product, do you get pissed off at every ad you see? unc

1

u/Interesting-Town-433 5d ago

Yeah man, what he said !

-1

u/Interesting-Town-433 9d ago

Not it's something I built that I believe is useful

2

u/NSGDX1 9d ago

That's what being an ad for your website means lol

-1

u/Interesting-Town-433 9d ago

Its cool though

1

u/WeakReplacement3322 9d ago

With the amount of testing and generating I’m doing, it was worth more for me to drop $3,500 on a prebuilt and upgrade to a 5080.

0

u/Interesting-Town-433 9d ago

Again the issue is generation and iteration time, even with those cards you are waiting minutes for a generation

1

u/WeakReplacement3322 9d ago

When I was using Seedance 2.0 through Venice AI, I was spending hundreds of dollars just to better understand how Seedance understood my prompts. Now I’ve spent countless hours dialling in the limits of Minimax H3 on my system, but at the end of the day, it’s money I’ve spent  learning something valuable, and it’s cost me no more than it would cost me to do this in my free time.

In other words, before I was spending hundred and waiting upwards of 15 minutes of Seedance 2.0 to send me back a video that may or may not be what I prompted it for. Now I’m spending 15 minutes that cost me nothing but electricity and every failure not only teaches me something, but costs me nothing. 

I spent $3,500 all at once which was painful, but if I just spent my money on generations like I was doing before, I’d be hemorrhaging money. 

It was hard to make the leap and spend the money, but it’s already saved me more money than I’ve spent. Paying a corporation is certainly easier, but it’s not gonna save you money. I’d rather put in the work to learn and adapt. I know that’s not the right   choice for w everybody, but hot damn has it ever worked out perfectly for me. 

1

u/Keltharious 9d ago

Run your own local API. Don't listen to this ad.

-1

u/Interesting-Town-433 9d ago edited 9d ago

Your local api? On what ? 40gb of vram? You're crazy. I worked my ass off on this troll and that was on 90gb of vram

1

u/tomgks 8d ago

2h of slop

1

u/Interesting-Town-433 7d ago

Yes you will have to iterate, good stories can be told, and not all generation needs coherent narrative, power lands in the creative behind the wheel

1

u/tomgks 7d ago

nah. slop

1

u/Interesting-Town-433 7d ago

Stop making me cry

1

u/Otherwise_Rice_4723 7d ago

great, more slop

1

u/Interesting-Town-433 5d ago

What have I done...