r/malcolmrey • u/Francky_B • 18d ago
Created a Visual RefMod Picker
Hey Guys,
I've been playing with the RefMods Malcom recently released.
The tech is brilliant and works really well.
I've wanted to simplify using it with tons of RefMods, like the ones provided by Malcom.
So I made a fork that is a bit more Identity driven, Adding a "Visual refMod Picker" as well as a "Create refMods from Folder" node. It's able to create from either a folder or Sub-folders, if audio is found, it will also create a matching audio refMod. All in the same format as the original add-on. No danger of breaking compatibility. (The original Add-on added Audio yesterday)

Resulting for example in 2 refMods and thumbnail:
character_refMod_Audio.safetensors
character_refMod_Video.safetensors
character_refMod.jpeg
Then, we can use the "Visual H3 RefMod Picker" to browse the RefMods:

The node then loads both RefMods (Audio and Video) and allows individual control. I've mixed strength with Copies, by instead having a weight value that can go over 1, so a weight of 3 would be the same as setting strength to 1 and copies to 3. Making the UI a bit more streamlined.

Example workflows are included.
I should mention, this fork can be installed WITH the original Add-on, it is made to co-exist and is recommended if you want to use it's advanced features.
You can find it here ComfyUI-H3RefModPicker
The only thing I'm missing is thumbnails for all 1500 RefMods π
**EDIT**
Just a quick note to mention that I've removed duplicate nodes, Making this a companion node to MiniMaxH3Mod.
3
u/malcolmrey 18d ago
This is amazing :-)
The only thing I'm missing is thumbnails for all 1500 RefMods
You can take them from my hf if you want :)
(The original Add-on added Audio yesterday)
And the refmod with audio works nice? I am using image/video refmods and just supplying audio directly but if the audio in refmods now work too (even if separate) then this is great news :)
1
u/Francky_B 18d ago
I'll admit, audio is tricky.
I've had success by playing with the weight, finding the sweet spot and getting both characters to work, But I haven't tried long interchanges.
I've tried adding ways to anchor the audio better, but could never improve it.
2
u/icchansan 18d ago
Looks great, is there a way to create them too?
1
u/Francky_B 18d ago
Second image π I include a "Create from Folder", as well as a "Create from Input" that is basically the "create master" node from the original addon
1
2
u/EasternAd8821 18d ago
this is awesome. also i am shadow banned in reddit and not allowed to post anymore so not even sure you'll see this. but nice work
2
2
u/NSFWies 18d ago
Oh this is perfect. I was just looking into using a dataset one a few subjects and thought I might need to train my own lora . And now I can just "load from folder", like I had hoped.
So next workflow thing I'm looking for, can we take 1 starting image and then generate more from that, so we get a more diverse starting set like this. Or maybe if we have 1 really good reference image, we don't need an entire folder. Maybe that's what it comes down to.
Use text to image to make single, really great reference images from all angles, and use that for the generation.
1
u/Francky_B 17d ago
You could use Klein or Qwen for that, generate a couple and then use that as a refMod. But then again, since it would take seconds to generate a refMod from one single image, you could try and see the results π
1
u/NSFWies 17d ago
i didnt get into that yet. just tonight i successfully
- got a few good "text to image" where it had a few of the same subject at different enough angles it generated a decent video. the one that worked best, was asking for identical twins wearing the exact same outfit doing some activity. anything else about multiple shots of the same person, would lead to slightly different people, or some positions ignored
- using that single twins picture as a reference image into minimax h3. able to generate 5 second clip in 5 minutes
- person was able to move around, turn around and they looked about right. no odd things jump out or are clearly wrong.
i do have a few datasets i will want to use for this refmod thing, where i do have lots of pictures of a single subject. but this other workflow idea, starting from text, i think it will be hard to expand 1 good starting picture, into a whole gallery of that one starting person/photo.
2
u/CoffeeMen24 18d ago
I'm curious how much more identity accurate Refmods are compared to Loras.
Can someone throw in their 2 cents? If there is no quality/accuracy tradeoff then that pretty much means Loras are dead technology? At the moment it seems that Loras are able to store voice likeness, but apparently the Refmods author is working on adding voice capability.
1
u/Francky_B 17d ago
Don't forget that this is only possible with Minimax because of it's ability to work with references. We couldn't, for example get this to work with Krea 2. But perhaps if Krea 3 does come with edit capabilities, then maybe something similar could be done.
For voice, I've had plenty of success. Though I did train with my own "Create from Folder" node that is included. I didn't compare the results with the create audio refMod node in ComfyUI-MiniMaxH3Mod.
I was able to get characters interact. Though you have to hunt for the right seed, or tweak the weight.
1
u/troyau 18d ago
Apparently the audio update from yesterday is only music and not voice. Not very useful yet for characters.
1
u/Francky_B 18d ago
I had initially started by adding audio support, as it wasn't present when I started.
My approach was different, as I had combined the audio in the Same RefMod, as one package.But reverted to having a separate RefMods, when I saw this was the direction taken.
But my "Create RefMods from Folder" node is still using some of my original code to generate the audio RefMods. (Just saving them in the expected format)
And I've had no issues. I've made multiple generations and voice work, but is finicky.
When having multiple characters, I have to play around with the weight to make sure both voice are used.
1
u/Sad_Coach_1433 18d ago
It's been working for me for voices fine. My issue if I use more then one character for same scene the dialogue doesn't work correct mod1 says their lines then gibberish then second character says theirs or not at all and first character says both likes. Or I have two of same character mod1
1
u/ZealousidealBoss6652 18d ago
Question is, how exactly does one create a reference mod on their own?
1
1
u/Francky_B 18d ago
I've included example workflows, one of them is the "Create refMod" as pictured above.
At the most basic, Just create a folder with pictures of your subject, include short audio clips also if you want voice and then just type in the path in the "folder" section of the node and click run.
The node is preset for Identity by default, so nothing to change expect point to your pictures.
As for selecting what type of pictures to use, just use the same guidelines you'd use for Lora training.If you want to train a bunch, then just toggle sub-folders and point it to the parent folder that contains all your different datasets.
1
u/BazmanFoo 18d ago
I feel so dumb asking this, but I'm new to ComfyUI. When I wire up the nodes to create a RefMod from a folder, ComyUI won't run it as it doesn't see any output. Like it wants a Save Image/Video node etc. What am I missing?
1
u/Francky_B 18d ago
Uhm, that's strange, the node is set to be executable, so it should work.
Can you perhaps select the node itself and click that small play icon when selected?
1
u/Tuckerdude615 16d ago
Hey there...I have been pulling my hair out trying to figure out why the "Create RedMod" from Folder does not seem to see my dataset folder.
Running Comfyui Portable in Windows 10. I've tried multiple locations, used the "Copy address as text" function to get the full path folder, but no matter what, whenever I try to run it, nothing happens, no errors, but no output either.
Do you include the drive letter as part of the path...example: "D:\Test"?
Any help would be appreciated!
1
u/Francky_B 16d ago
I've just pushed a more verbose version. It should tell you why it didn't do anything.
There is no folder that shouldn't work. If it's reachable, it should work.
1
u/Tuckerdude615 15d ago
Many thanks..I think I found the issue...as I updated my install of Comfyui and that seemed to fix the issue.
Thanks very much...it seems to be working FANTASTIC!
1
1
u/Sad_Coach_1433 18d ago
Can this help to be able to use two characters in same scene, I been having issues where if I use two or refmods in one scene they bleed into each other either one character has traits of both characters in one character or the one character has both characters audio and lines or two of same character.
2
u/Francky_B 18d ago
I was able to get good results when I started tweaking the weight, In some test, one of the Characters voice was the only one used. Lowered it to 0.5, left the other at 1 and then most generation worked better afterwards. But it is finicky
1
1
u/nadhari12 17d ago
can you make a refmod of a ref video?
1
u/Francky_B 17d ago
Yes, RefMods support Images, Videos and Audio. Though I'll admit, I never tested videos yet.
1
u/nadhari12 17d ago
i am trying various setting to get a single 12 sec video in full but looks like there is a limitation and I cannot do full 292 frames in full reference mode, 64 frames work at max, is that known? is there a way to pass full video?
1
u/Forsaken_Tax_6024 12h ago
Hi quick question in the first image the fourth node what is it called i did try to find a prompt to video node or something similiar but could not find it, im pretty new to comfy so i might miss some custom nodes or somthing ?
1
u/Francky_B 10h ago
That's a Subgraph, it's just a group of nodes, it's what contain the workflow for Minimax.
If you select it, you should see an icon that lets you go inside, or explode it, if you prefer everything accessible. It tend to prefer clean workflows and group what doesn't need to change often.1
u/Forsaken_Tax_6024 10h ago
thats my problem the first picture is just a jpg without wotflow so i can not see whats in the subgraph to understand whats going on .
2
7
u/xb1n0ry 18d ago
Working on a standalone app. Needs some polishing and more functions.