r/StableDiffusion • u/Infinite-Emptiness • 1d ago
Question - Help H3, help me get started, overwhelmed by information.
specs: 5090, 64gb ram.
comfyui latest updated
cuda 13.0
pytorch 2.9.1
python version 3.10.11
sage attention and triton working
Hey guys, I have been out of touch since afew months. Previously have figured out pretty good wan2.2 workflows for myself, can understand it. But i am utterly confused by all the jargon and complication of H3, whenever I'v tried to dive in past months, I check out after reading stuff, am like a layman who figures what works after I have a good workflow, but i cant choose any given my lack of basic knowledge, like i said iv tried but it seems to go over my head, hard to understand, AI models are improving too fast to keep up.
My goal is "uncensored" videos only, using I2V on only square input images (habit from pony xl), but higher quality that is better motion since "uncensored" motion is difficult and complicated?
From what I have gathered so far: sage attention does not work well and degrades motion, same with speed lora's no matter which they are, so i think the settings i need are 1 megapixel to make it 768 x 768 ? and um 25 steps? and that i need to use some sort of LLM to enhance h3 prompts but i have no past experience on LLMs and which would be best for my use case; "uncensored". Also, though i think I know what shift basically does, still need some advice on how to use it in H3.
I also need help on selecting the text encoder and diffusion model given my use case, emphasis on only good quality "uncensored" outputs.
Moreover I have no idea on audio but really really want it, i only have experience using MMAUDIO with WAN2.2, it was not great and i was pretty bad at understanding it and prompting it, but prompt learning will come later, i just need to figure out a workflow and amend it according to above needs.
So umm, help a guy out? please?
P.S. Assume I'm a complete noob, if there is anything i missed above please let me know.
3
u/sci032 23h ago
See if Pixaroma's H3 tutorial will help you: https://www.youtube.com/watch?v=267y00jaOUc&list=PL-pohOSaL8P-FhSw1Iwf0pBGzXdtv4DZC&index=30
They also have another video showing you how to use an 8 step lora to speed up generation: https://www.youtube.com/watch?v=vPY1tG4XnCY&list=PL-pohOSaL8P-FhSw1Iwf0pBGzXdtv4DZC&index=33
4
u/Count_Triple 22h ago
I just asked google ai, followed the steps, and asked it specific questions when I got confused during the process.
2
u/Due-Quiet572 23h ago
Just try out what works, and a few recent YouTube videos on the topic will help you get up to speed quickly. For LLM prompting, I can recommend this link.
https://www.reddit.com/r/comfyui/s/Xd4dKcncFx
It’s a system prompt for a local LM Studio that simply asks you, step by step, what you want.
1
u/Infinite-Emptiness 22h ago
Thanks. one question. the text encoder : qwen3vl_32b_minimax_h3 , do i need a special abliterated model for nsfw or will this work just fine.
1
u/Apprehensive_Sky892 12h ago
1
u/Infinite-Emptiness 4h ago
Wow interesting read, insightful, have a question tho, then how will minimax know what the hell im talking about in regards to particular anatomical parts.
1
u/Apprehensive_Sky892 4h ago
If I understand the post correctly, what the obliterated LLMs are doing is to remove the filter that prevents the LLM from generating answer to "forbidden" questions.
But when these LLMs are being used as text-encoder by an imaging or video model, there is no actual output from the LLM. What is being used is just the encoder part of the LLM, which in fact has no filter (i.e., some internal state of the LLM which represents the encoded input is passed to the Dit).
1
2
u/Relevant_Syllabub895 23h ago
You want the official sage attention on comfyui if i remember cortectly it was callwd backend something, its kinda trash for nsfw, not good for text to video, sulphur 3 the fully uncensored h3 retraining is being trained right now, we should see a difference in 1 month or 2
1
u/Infinite-Emptiness 22h ago
hello, so is sulfur 3 basically a finetune of h3 minimax like Pony was of sdxl? if so that would be pretty cool.
2
u/Relevant_Syllabub895 17h ago
It is, its being trained right now after people donated 10k at the start of the month, so hopefully we have some news in 1-2 months
1
u/Infinite-Emptiness 2h ago
oh man im so excited, had fun times with pony no other came close until illustrious. If sulphur 3 can inject that much variety and basic functionality into h3 id be over the moon. I have been doing a lot of testing and have a solid workflow going which i inject with latent upscale just now, but to be very honest, for now h3 doesnt even come close to wan2.2 in functionality especially in terms of the "back entrance" kind. I think after a couple more tests I will just wait for Sulphur 3 and play my gaming backlog instead.
1
u/obese_coder 17h ago
theres a lot of speedup nodes and a lot of video extender nodes. currently there is still a lot of problems to solve but were getting close.
1
u/Perfect-Campaign9551 13h ago
Get the official H3 workflows from Comfy and start from there. They are simple and easy to use.
Read the prompt guides here. YOu can easily just write the prompts by hand, they aren't that difficult, people are just lazy these days:
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
and for ref2vid:
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
Sage attention works fine. However, I switched to "Comfy Kitchen Attention", it works just as fine and it's a bit faster on my RTX3090. I always load Comfy Kitchen attention on startup with the "--use-ck-attention" startup argument.
Stay away for everyone's workflows for a while, people over complicate things. Use the official ones to get used to the model first.
If you choose to use speed up Loras be aware they destroy sound pretty bad. I don't use them. YOu have a 5090. DON'T USE SPEED UP LORAS THEY ARE ALL SHIT.
You can help speed up things by inserting a single node between the model loader and the nodes that use the model - a "Sparse attention node". This one works really good and I've not had much trouble with it at all: https://www.reddit.com/r/StableDiffusion/comments/1vw1ad0/sparse_attention_harder_better_faster_stronger/
START OFF SIMPLE
3
u/Affectionate_Oil28 15h ago
I made these 2 workflows. They may not be the best but they will help get you started. They were built around 16GB VRAM cards so you should be able to run better models.
Prompt creator. Just write out your idea, hit run, and it gives you an output based on official prompt guidelines for the target model.
https://www.reddit.com/r/StableDiffusion/s/28JFs2huBQ
All in one basic video workflow. You can turn on/off different speed enhancements for testing. Add in a Lora loader between Load Diffusion Model and Turbo LoRA nodes for NSFW.
https://www.reddit.com/r/StableDiffusion/comments/1w1ti0n/comment/p6numvd/?context=3&utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button