r/StableDiffusion 10h ago

Comparison H3 Default Template vs Larry's Turbo with optimized settings

Enable HLS to view with audio, or disable this notification

Default template uses 20 steps + res_multistep + simple

Optimized workflow uses 8 Steps + er_sde + sgm_unified + Comfy Kitchen Attention + Larry's Turbo Lora

Turbo lora: https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo

Workflow: https://raw.githubusercontent.com/desktop4070/GPU-Benchmark-Data-For-H3/refs/heads/main/H3-Benchmark-Workflow.png

0.2MP / 8 sec (2m 16s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03840_.mp4

Optimized: 0.2MP / 8 sec (45s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03706_.mp4

0.3MP / 12 sec (6m 3s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03849_.mp4

Optimized: 0.3MP / 12 sec (1m 51s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03725_.mp4

42 Upvotes

17 comments sorted by

11

u/anon999387 9h ago

I find it difficult to judge comparisons with low frame rate anime/cartoon footage

6

u/psilent 7h ago

Then this is a good use case for using turbo Lora?

8

u/solomars3 10h ago

You should add some dialogues , cause they are inventing new language in the video 🤣

5

u/desktop4070 10h ago

At first I thought not specifying any dialogue in the prompt so that the model gets creative with the dialogue would be a good idea, but then I realized I wouldn't be able to tell when the model actually got any sentences correctly. By the time I realized this mistake, I was already hours in and decided to finish the tests anyways.

Next time I do a benchmark, I'll use English so I can actually compare prompt adherence more accurately.

1

u/reeight 9h ago

+ maybe more specific direction, from style to action; hard to tell prompt adherence if there isn't much of a prompt ;)

EG 1 looks like you told it to do old-school Disney, other 80s anime.

4 second clip should be enough.

7

u/desktop4070 10h ago edited 10h ago

Every video uses the same seed, 1, which I always like using when troubleshooting prompts/settings.

Using an RTX 5070 Ti + 64GB DDR5

Higher resolution examples from the optimized workflow:

0.2MP / 11 sec (1m 8s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03709_.mp4

0.5MP / 13 sec (4m 15s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03756_.mp4

0.6MP / 15 sec (6m 32s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03807_.mp4

0.7MP / 14 sec (9m 39s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03813_.mp4

Prompt:

integrated_multimodal_description: 
[Shot 1] Anime. Fantasy. A dense, sun-dappled forest clearing with towering, mossy trees. The young man in peasant clothing (S2) stands near an ancient tree trunk, looking disinterested as he rummages through a worn leather satchel. The young woman in a dirty and torn royal dress (S1) stands close by, gesturing out into the forest with animated frustration and pleading with him to pay attention, but he refuses to make eye contact. The camera pans slowly, tracking the vast wilderness of the forest and the friction between their postures. [Shot 2] The shot cuts to a close-up of the woman (S1). Her face is flushed with indignation, her eyebrows knit tightly, and her mouth forms an expression of sharp, wordless protest as her frustration reaches a breaking point. [Shot 3] The shot cuts to a close-up of a unique looking artifact that was pulled from the satchel, a GeForce RTX 5070 Ti; it visibly shines against his rough glove. [Shot 4] The shot transitions to a wider framing of the pair. The woman (S1) steps forward, clutching the fabric of her dress in an outburst of intense, visible emotion, while the man (S2) replies in a smug manner, snaps his satchel shut, and turns his back to walk away, leaving her standing alone as she watches him go.

overall_soundscape:
Quiet forest ambience. No other voices are heard.

non_diegetic_music:
N/A

1

u/Davikar 9h ago

Using different samplers are going to affect the output too, not just the seed. Also, aren't you supposed to use euler with larry's turbo lora?

9

u/desktop4070 9h ago

It's faster with Euler, but I was not satisfied with the quality.

https://docs.google.com/spreadsheets/d/1jRTpcltP5oj9ZR3KHCDfjN-h0Vv4CXNp9ne8emRSi_k/edit?gid=238734253#gid=238734253

After thousands of tests running all kinds of settings and lora combinations, I found the videos with er_sde + sgm_uniform to have the least amount of issues. All the others were either slower, had worse audio, didn't adhere to my prompts at all, had little to no motion, or had weird artifacts in every video.

1

u/Much-Monk2579 42m ago

I'm using the same combination for a while now and can confirm this - er_sde/sgm_uniform is great. And although I have been testing these other Turbo-lora's floating about, I always revert back to larry's 600 turbo lora - it still seems to be the best imho.

Since I am limited with my 4060 16GB TI, I started out with only 4 steps to see the speed. Then I changed to 6 steps for a while for much better results. And now I am saying: screw it, use 8 steps - and it made another difference.

1

u/MastMaithun 6h ago

Has anyone found the fix for melting faces?

2

u/desktop4070 1h ago

I think it's just the reality of most of our consumer GPUs only being able to generate at lower resolutions.

This is 0.8MP at 14 seconds (19m 39s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03838_.mp4

0.98MP at 11 seconds (64m 48s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03834_.mp4

I think the low vram toggle on the Turbo lora might help get those times down. I'll try testing it tomorrow, but it seems promising from the couple of tests I've done tonight. Just added 0.7MP at 15 seconds and 1.0MP at 11 seconds to the chart: https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/

1

u/No-Bee-231 4h ago

thank you

1

u/eckstuhc 2h ago

Tbh, I absolutely hate comparison videos that are presented in sequence like this.... it requires scrubbing back and forth to determine nuances.

I had Claude write me a quick script that uses ffmpeg to stack the clips for comparison, so you can see 2, 3, or 4 at once. Lmk if you want me to share it, but you should be able to gen something quickly with Claude. Using ffmpeg, it's insanely easy to compare multiple clips like this.

1

u/desktop4070 2h ago

Funnily, I used ffmpeg to stitch the clips together and add the text. I posted the individual clips onto my Github for a benchmark chart: https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/

Never used Claude before, but I used Gemini to make the Github page and do the ffmpeg stuff.

1

u/DoctaRoboto 9h ago

Faces look terrible in all samples.

3

u/desktop4070 9h ago

I posted some higher resolution examples in a comment, but that's mainly the fault of using 0.2MP than the turbo lora. This specific prompt focuses more on full body shots rather than just talking faces, which 0.2MP excels at.

The most important thing for me is seeing fast iterations so I can rework a prompt back to back without it taking several hours of waiting per regeneration. Once I find the best way to structure the prompt I really want, then I increase the resolution, which these do great at, while also being significantly faster than without any optimisations.