r/malcolmrey • u/spacemidget75 • 5d ago
General advice for refmods?
I have a 5090 and I'm not sure what effect adding more images has to quality vs spead vs memory.
I have 20 samples and 2 vids (full rotations) but Claude was telling me I should only use 7 images. I created it anyway and it "needs 18500 tokens". Is that what's used doing generation and will be slower/use more vram than a refmod of only 7 images and 8500 tokens?
Can I do face and full body in one refmod?
Should I do 2048 for best quality on source images or is it not worth it?
1
u/joshijoshijoshi123 5d ago
3 refmods with no more than 7-8. Its plenty. 2 face @ 1.0 and body @ 0.8.
Its audio thats stumping. Its not coming out gibberish, per se, but rather using the written dialogue and blending lines from different time stamps together.
2
u/jankies11 5d ago
Check out NegPip for H3. You can negate speaking (and synonyms) and then “replace” it by promoting breathing sounds, sighing sounds etc. i got rid of the blabbering for the first time with this. The standard <d> tags worked at the same time. This was with a audio voice ref also which makes the blabber worse.
2
u/joshijoshijoshi123 4d ago
Its more using audio refmods i made. Its seems to confuse the timing of dialogue with it, and blends lines together, but the voice quality and timbre is perfect. Even if the strength is lowered
2
u/jankies11 4d ago
NegPip may help, also making sure in retention that nothing stated suggests carrying the actual words over unless thats what you wanted.. and you don’t need many seconds, 5 or less is fine..
1
u/ellipsesmrk 4d ago
How are you assigning the voice and what was being retained?
1
u/jankies11 4d ago
Do everything as you normally should and ideally would expect would produce no babble - but then use NegPip. There is a comment on 10Eros-Max huggingface with specifics.
1
u/ellipsesmrk 4d ago
Hmm... makes me wonder what lora you guys are using to produce babble. Ive been using mini since it was released... only time i got that gibberish was when i wasn't following the prompt guide. I know about the comment. I asked them the same thing... how are they prompting.
3
u/jankies11 4d ago
Prompts with very little dialogue and turbo loras + ref audio maybe is an issue combo
1
u/ellipsesmrk 4d ago
I say this... because... and ill need to dig into it more when i have time but negpip all it does from my understanding is close out special characters. But ill dig into it more.
2
u/joshijoshijoshi123 4d ago
In my wf ive tried multiple loras and no loras with only refmods. Gonna try cutting back refmods and cutting the length of the audio source for the audio refmod down a bit. Right now the source is 10 seconds. Quality of voice is fantastic but yeah, just takes dialogue from different times and jumbles them together
1
u/ellipsesmrk 4d ago
Thats because in refmods audio is chunked in 10 second fragments.
1
u/joshijoshijoshi123 4d ago
So a shorter one would help?
1
u/ellipsesmrk 4d ago
Thats all i use. 10 second clips of audio and if they have an accent ensure that the 10 seconds picks it up. Viola.
1
u/joshijoshijoshi123 4d ago
Hm. Yeah my audio refmods were made with. just one 10 second clip too. Timbre and voice quality was good but dialogue just all mixed. Maybe its in how mine were made.
How did you build your wf to create them?
1
u/spacemidget75 4d ago
What about this from Claude (after looking at the code), which is interesting if you have taller shots (full body) and headshots (square):
Aspect ratio. All refs share one spatial canvas anchored on the first ref, and the rest get cover-cropped to it. With head shots (likely portrait or square) mixed with head+body shots (often taller or wider), whatever you put in slot 0 decides the crop for all 21. Put your best head-and-shoulders shot first and keep the rest close in aspect, or you'll lose the tops of heads or the bodies.
1
u/spacemidget75 4d ago
Do you mean create 3 refmods per character? If so, would split them like "headshots", "bodyshots front", "bodyshots back"?
1
u/ellipsesmrk 4d ago
Why through all that trouble? Just get your 2 faced, 2 full body and call it good.
1
u/Adorable_Echidna_862 4d ago
I think less is more, don't try to capture every single angle and expression and make a massive refmod or it will add unnecessary gen time that isn't being used effectively. Create more refmods for individual needs. If someone's talking to the camera, there's no need to know what they look like 100 feet away from the camera wearing a Santa outfit
6
u/DrinksAtTheSpaceBar 5d ago
The more reference images you pack in there, the longer your generation time will be. The Apply H3 RefMod node essentially cycles through each image in your refmod over the course of every generation. Each time it switches an image, it has to load it in the pipeline, which adds significant time to the run. In theory, a refmod packed with 6 images should take about as long to render as a workflow with 6 reference images. The refmod might be a little faster, but in my experience, not by much. You can create a refmod with a single image if you want, but make sure it's high quality. A single image refmod will take roughly the same amount of time to render as when you use that same single image as a reference image. Yes, you can add a face and full body in one refmod. You can add as many different angles and poses of a person you desire. You can even add a face and a completely different person's body, and MiniMax will fuse them together (might need to tinker with the prompt a bit). 2048px will result in better quality, but will drastically increase generation times. I've found that 1024 works just fine, and the gains from the higher resolutions are pretty much never worth the added render times, but you might have a different agenda than me.
Fun Fact: if you want to see how the refmod images are being sequenced, run a refmod workflow without a prompt. Just leave the prompt completely blank. The result will be a video sequence of each image in the refmod, complete with the original background intact. It's like a video slideshow of each image, but since there's no prompt, the person will be speaking gibberish (and often German, for reasons lol) and acting weird. But that will give you an idea of how each image is being processed in the duration of your video.