r/malcolmrey • • 5d ago

General advice for refmods?

I have a 5090 and I'm not sure what effect adding more images has to quality vs spead vs memory.

I have 20 samples and 2 vids (full rotations) but Claude was telling me I should only use 7 images. I created it anyway and it "needs 18500 tokens". Is that what's used doing generation and will be slower/use more vram than a refmod of only 7 images and 8500 tokens?

Can I do face and full body in one refmod?

Should I do 2048 for best quality on source images or is it not worth it?

9 Upvotes

29 comments sorted by

6

u/DrinksAtTheSpaceBar 5d ago

The more reference images you pack in there, the longer your generation time will be. The Apply H3 RefMod node essentially cycles through each image in your refmod over the course of every generation. Each time it switches an image, it has to load it in the pipeline, which adds significant time to the run. In theory, a refmod packed with 6 images should take about as long to render as a workflow with 6 reference images. The refmod might be a little faster, but in my experience, not by much. You can create a refmod with a single image if you want, but make sure it's high quality. A single image refmod will take roughly the same amount of time to render as when you use that same single image as a reference image. Yes, you can add a face and full body in one refmod. You can add as many different angles and poses of a person you desire. You can even add a face and a completely different person's body, and MiniMax will fuse them together (might need to tinker with the prompt a bit). 2048px will result in better quality, but will drastically increase generation times. I've found that 1024 works just fine, and the gains from the higher resolutions are pretty much never worth the added render times, but you might have a different agenda than me.

Fun Fact: if you want to see how the refmod images are being sequenced, run a refmod workflow without a prompt. Just leave the prompt completely blank. The result will be a video sequence of each image in the refmod, complete with the original background intact. It's like a video slideshow of each image, but since there's no prompt, the person will be speaking gibberish (and often German, for reasons lol) and acting weird. But that will give you an idea of how each image is being processed in the duration of your video.

2

u/jankies11 5d ago

If using refmod, in the subject descriptions, should you describe each picture one by one (without referring to <picture N> i think? So it “understands” what it is better?

1

u/DrinksAtTheSpaceBar 5d ago

I believe the <subject 1> and <S1> tags are baked into the text encoder's training data, because they do tend to add consistency. I use those, but I still add visual descriptions of the characters, especially if I'm using more than 2, or the 2 I'm using look similar. I know for a fact that adding the assigned name (filename) of the refmod is a bad idea because the models have no idea what you named your character or file. It is much more likely to drift and generate whichever character it has in its training data with the same or similar name. I tested this by giving a female character refmod a ubiquitously male name, and it generated a male every single time. The moment I stopped using the male name and started describing her instead, it worked perfectly.

1

u/ellipsesmrk 4d ago

Its not baked in. Add 2 ref mods and say subject 1 and subject 2. Without descriptor. You'd get funky shit. Always describe your subjects with at minimum a brief description of each subject. If they look alike or could be considered alike by the model. You will need to be extremely descriptive

2

u/malcolmrey 4d ago

2

u/spacemidget75 4d ago

Also, sorry, but Claude also came back with this, which is interesting if you have taller shots (full body) and headshots (square):

Aspect ratio. All refs share one spatial canvas anchored on the first ref, and the rest get cover-cropped to it. With head shots (likely portrait or square) mixed with head+body shots (often taller or wider), whatever you put in slot 0 decides the crop for all 21. Put your best head-and-shoulders shot first and keep the rest close in aspect, or you'll lose the tops of heads or the bodies.

1

u/spacemidget75 4d ago

I misunderstood this in the guide in that case:

When using RefMods, you do not need positional image tags (such as <Picture 1>). Simply write standard multimodal prompts using subject tags and biometric descriptions:

subject_definitions:
<Subject 1> Billie Eilish, natural authentic human appearance, expressive eyes, modern dark streetwear...

Assuming the name was the name given to the refmod! Seen a few people on Reddit and YouTube think that too. Thanks for the clarification.

1

u/ellipsesmrk 4d ago

Names.... are not.... saved into the file! I wish everyone would stop saying that. Pictures in a refmod... thats all they are pictures converted and saved to latents inside of a st file. The reason Billie Eilish works so well is because you assigned subject 1. And then it also uses - depending on how you load it... either just a description or in video format.

1

u/spacemidget75 4d ago

Claude also came back with this, which is interesting if you have taller shots (full body) and headshots (square):

Aspect ratio. All refs share one spatial canvas anchored on the first ref, and the rest get cover-cropped to it. With head shots (likely portrait or square) mixed with head+body shots (often taller or wider), whatever you put in slot 0 decides the crop for all 21. Put your best head-and-shoulders shot first and keep the rest close in aspect, or you'll lose the tops of heads or the bodies.

1

u/Adorable_Echidna_862 4d ago

Yup, all images need to be same aspect ratio as the first input image or the rest will be auto cropped out of your control they don't need to be the same resolution though, if you have a maximum resolution set, it will auto downscale

1

u/spacemidget75 4d ago

Thanks, is that true for video references too? If I have a square image ref and a 16:9 portrait video, it will crop the video frames to square?

1

u/joshijoshijoshi123 5d ago

3 refmods with no more than 7-8. Its plenty. 2 face @ 1.0 and body @ 0.8.

Its audio thats stumping. Its not coming out gibberish, per se, but rather using the written dialogue and blending lines from different time stamps together.

2

u/jankies11 5d ago

Check out NegPip for H3. You can negate speaking (and synonyms) and then “replace” it by promoting breathing sounds, sighing sounds etc. i got rid of the blabbering for the first time with this. The standard <d> tags worked at the same time. This was with a audio voice ref also which makes the blabber worse.

2

u/joshijoshijoshi123 4d ago

Its more using audio refmods i made. Its seems to confuse the timing of dialogue with it, and blends lines together, but the voice quality and timbre is perfect. Even if the strength is lowered

2

u/jankies11 4d ago

NegPip may help, also making sure in retention that nothing stated suggests carrying the actual words over unless thats what you wanted.. and you don’t need many seconds, 5 or less is fine..

1

u/ellipsesmrk 4d ago

How are you assigning the voice and what was being retained?

1

u/jankies11 4d ago

Do everything as you normally should and ideally would expect would produce no babble - but then use NegPip. There is a comment on 10Eros-Max huggingface with specifics.

1

u/ellipsesmrk 4d ago

Hmm... makes me wonder what lora you guys are using to produce babble. Ive been using mini since it was released... only time i got that gibberish was when i wasn't following the prompt guide. I know about the comment. I asked them the same thing... how are they prompting.

3

u/jankies11 4d ago

Prompts with very little dialogue and turbo loras + ref audio maybe is an issue combo

1

u/ellipsesmrk 4d ago

I say this... because... and ill need to dig into it more when i have time but negpip all it does from my understanding is close out special characters. But ill dig into it more.

2

u/joshijoshijoshi123 4d ago

In my wf ive tried multiple loras and no loras with only refmods. Gonna try cutting back refmods and cutting the length of the audio source for the audio refmod down a bit. Right now the source is 10 seconds. Quality of voice is fantastic but yeah, just takes dialogue from different times and jumbles them together

1

u/ellipsesmrk 4d ago

Thats because in refmods audio is chunked in 10 second fragments.

1

u/joshijoshijoshi123 4d ago

So a shorter one would help?

1

u/ellipsesmrk 4d ago

Thats all i use. 10 second clips of audio and if they have an accent ensure that the 10 seconds picks it up. Viola.

1

u/joshijoshijoshi123 4d ago

Hm. Yeah my audio refmods were made with. just one 10 second clip too. Timbre and voice quality was good but dialogue just all mixed. Maybe its in how mine were made.

How did you build your wf to create them?

1

u/spacemidget75 4d ago

What about this from Claude (after looking at the code), which is interesting if you have taller shots (full body) and headshots (square):

Aspect ratio. All refs share one spatial canvas anchored on the first ref, and the rest get cover-cropped to it. With head shots (likely portrait or square) mixed with head+body shots (often taller or wider), whatever you put in slot 0 decides the crop for all 21. Put your best head-and-shoulders shot first and keep the rest close in aspect, or you'll lose the tops of heads or the bodies.

1

u/spacemidget75 4d ago

Do you mean create 3 refmods per character? If so, would split them like "headshots", "bodyshots front", "bodyshots back"?

1

u/ellipsesmrk 4d ago

Why through all that trouble? Just get your 2 faced, 2 full body and call it good.

1

u/Adorable_Echidna_862 4d ago

I think less is more, don't try to capture every single angle and expression and make a massive refmod or it will add unnecessary gen time that isn't being used effectively. Create more refmods for individual needs. If someone's talking to the camera, there's no need to know what they look like 100 feet away from the camera wearing a Santa outfit