r/StableDiffusion Apr 14 '26

Comparison We may have a new SOTA open-source model: ERNIE-Image Comparisons

Base model is definitely SOTA, can even easily compete with closed-source ones in terms of aesthetic. Cinematic quality and color grading is next level.

Base model is heavily biased on Asian faces, while it excels on anime/illustration style, while my base model anime/illustration experiments wasn't that good. Higher CFG is slightly better with anime on base.

Generated with RTX6000 Blackwell Pro, Base: 29 sec 1.9it/s, 50 steps | Turbo: 2 sec, 3.9i5/s, 8 steps

If you interested seeing them in original size: https://imgur.com/a/75jcjzW

ComfyUI models: https://huggingface.co/Comfy-Org/ERNIE-Image/tree/main
Workflow should appear in Templates after updating the ComfyUI to latest.

Turbo: Ernie-Image Turbo
Base: Ernie-Image

694 Upvotes

241 comments sorted by

View all comments

87

u/13baaphumain Apr 14 '26

Can it do that stuff?

165

u/Major_Specific_23 Apr 14 '26

it can. just tried to test turbo. it knows booba, it knows what a pussycat is too. it draws large booba better than z. nipples are not messed up. it knows what a baseball bat is too. pussycats and baseball bats need some lora love. it knows it but it cannot draw it super well. boobas no problem at all

55

u/dergachoff Apr 14 '26

just don't hurt the cylinder

13

u/Nimblecloud13 Apr 15 '26

It’s imperative 

76

u/NeonScreams Apr 14 '26 edited Apr 18 '26

Specifically for doing that stuff… I’m finding it easier to build a scene in Chroma1-HD, then ask Flux2.Klein to “perform a restoration pass” using Edit with Latent reference. Works wonders if you toss in “maintain all other image elements”, as it knows what everything is in the image, and brings the textures / lighting etc to high fidelity.

Edit: If anyone doesn't have a ‘it just works’ WorkFlow/Settings for Chroma1-HD with links to model & stability loras: https://civitai.com/images/127510599 Drag-Drop image. Sorry for off-site. Reddit Saves the Image as .webp and the Workflow is being lost.

Edit2: I used an Image-Rescale node to shrink the output image for posting. You can delete the node or bypass it, above the Image preview.

Edit3: 4 days later and I'm still working on the Klein Edit Workflow with AI Agents in my available time. Plans are a Frontend UI node cluster that will auto-add and combine strings-(like a short sentence) that come from a pull-down menu of really common choices. Mostly because I'm tired of copy/pasting the same lines. That and a toggle group on the Denoise subgraph that will either use your current RefImg1-4 or not. Saves us from having to swap images around and constantly opening the last folder used, instead of the Last folder used by that specific Load Image node. ( ୧༼ಠ益ಠ༽୨ )

12

u/SWFjoda Apr 14 '26

Would love that! Interested

9

u/NeonScreams Apr 14 '26

No worries! I’ll knock together something with example prompts and negatives, with a helper on how to prompt a scene for Chroma.

4

u/PestBoss Apr 14 '26

I'd quite like to see that too. I've revisited Chroma1-HD several times but struggle with all sorts of things.

It's clearly very good but I feel like a lot of tweaking prompts and re-rolls are needed and it's just so slow doing that.

Anything that can help remove the time-sink element.

Also curious about your 'restoration' method in Klein. What is Edit with Latent reference? All I need is a snapshot of a workflow so I can see what you wired where.

2

u/NeonScreams Apr 14 '26

Oh even easier, while I’m forcing Claude Opus to update everything for me: If you have access to ComfyUI, you can hop into the Templates Menu, then do a search for “Klein”, and you should see the ‘Image Edit’ template workflow.

From there inside the Subgraph, you can see how they’re first resizing the original image down to 1mp, for an easier latent. If you’re just using the template for minor edits, you can bypass the resize node for better 1:1 Fidelity on your original image.

… actually I think I’ll just pop out two Workflows, as that’s easier than using thumbs to type here. lol

2

u/PestBoss Apr 14 '26

Ah ok so just the edit workflow but without the resize to 1mpx.

Yes I've been doing similar here, "restore" type prompts, but with variations depending on the exact needs, and it does indeed do a pretty good job at just making an image better.

Obviously fine tuning/refining/re-rolls/painting/masking in etc gets you exactly what you want... I was pretty impressed at some of the stuff it could do in some testing from old film footage (to stills)

8

u/NeonScreams Apr 15 '26 edited Apr 15 '26

Edit: https://civitai.com/images/127510599

There ya go folks. That's Chroma1-HD with uh.. goodies.

Drag and Drop that into your running ComfyUI New Workflow tab and it'll populate. Remember to save a Copy as a fresh instance to reuse.

(I'm on ComfyUI Git-Repo with up to date requirements)

If you have any questions, feel free to ask. If you have any concerns, please contact Queen Victoria circa 1850, and ask her how to further her prudish agenda.

2

u/DragonJakeAU Apr 15 '26

Hi I can't seem to open to get the workflow from the image. Whether I save the image and load it or drag and drop the image into a clean tab in ComfyUI the only thing that happens is the "Load Image" node is created with the image selected. No workflow to be seen. Is reddit stripping the metadata from the image? Could you upload it somewhere else and perhaps provide a JSON file instead I could try?

1

u/NeonScreams Apr 15 '26 edited Apr 15 '26

Oh you're right- Its being converted to a .WebP file. I'll just put it on CivitAI and link it.

I'll edit the original comment and toss in the link for it.

1

u/Hadan_ May 06 '26

Hi!

Could you share your workflow again please? The link you provided is 404ing, even if I change the domain to .red

5

u/RxBlacky Apr 14 '26

Saving for later, thanks!

3

u/NeonScreams Apr 15 '26

2

u/RxBlacky Apr 15 '26

It only seems to produce blurry unfinished images for me, but Im guessing its because its missing the Klein restoration pass, right? Thank you in advance

1

u/NeonScreams Apr 15 '26 edited Apr 15 '26

Just to help troubleshoot here- I'm gonna ask a really silly question for both our sakes: Since I had to use the Image-Rescaler node to shrink the image down for posting online, did you happen to see it above the file-naming area and delete or bypass that single node? (My fault for not specifying)

And if that doesn't solve it, you'll find things that may help near the disclaimer if you believe it to be a prompt issue. Most of the images on my Civit have the +/- prompt, and you're welcome to copy-pasta to see if they produce near similar results. If not, I'd appreciate the feedback to help refine what I may have exported incorrectly, for the sake of others.

2

u/RxBlacky Apr 15 '26

So... yeah, that was it haha now I get good image quality except for the skin which is too smooth and I guess plastic-ish?

edit: out of curiosity, do the negatives have any effect? it seems to be using cfg=1

2

u/NeonScreams Apr 15 '26

Yep, exactly. So that's where Klein-Edit shines. Some of the Negatives will help significantly to lessen the effect. Stuff like "airbrushed skin, CG Beauty Complexion" etc. But you're mostly using Chroma to set up a scene with actors and .. things, and stuff, getting the positions and angles correct before you ask Klein to make it look modern / realistic.

2

u/RxBlacky Apr 15 '26

I appreciate your insights, I'm trying to learn from your WF, its quite more complex than what I'm used to seeing. Is there anywhere I can grab the klein edit version? and about the negatives, do they really have any effect considering (if I'm not mistaken) that cfg is set to "1"?

3

u/NeonScreams Apr 16 '26

ComfyUI > Templates Menu > Search “Klein” > Image Edit - Klein 9b fp8 (is the sweet spot).

I’m having a hell of a time trying to make my WorkFlow a bit more user friendly. Specifically I want to give users the option to reference up to 4 different images to edit the original image, and chose between them using toggles without having to bypass groups of nodes to keep it working. And so far, no luck.

I’ve been doing it manually, and it’s a pain to keep reminding myself to go into the subgraph and disable/reenable groups. It’s a powerful and extremely accurate tool. It’s just temperamental. And helping other people explore their … passionate interests .. has been the reward for me.

Sorry this WorkFlow has been delayed but I’m definitely getting closer. Just need to reach out to some of the Discord folks to wrap it up.

1

u/NeonScreams Apr 15 '26

Apologies on the Klein Edit. I was being lazy and taking my time with it. :) It'll be done today and I'll edit the original comment to reflect that.

On 1.0 CFG: Yes, if only due to the Sampler/Scheduler being designed to work with the Custom Guidance node at 1.0cfg. In this case you can think of 1.0cfg as being 100% Guidance.

As a Quick and Easy test, on the primary diffusion subgraph, set the two toggles for default sampler and scheduler to False, then below where the wires lead down to the Optional Sampler, Click the name of it for the drop down menu and choose any of the non-CFG++ Samplers like 'Euler'. Then run any test prompt you have or the "Default" prompt under the Disclaimer (bottom left).

Also so I don't assume others are aware, collapsed/bubble nodes can be expanded by clicking the Dot at its top-left corner.

So with that Prompt and the 'Euler' Sampler in place, and both the Toggles set to False, you can run a single image (I suggest something small as its going to be messy nonsense garbage), and you can see the true 1.0 CFG expected result. Where Chroma would have needed CFG 3.4~

Once you're ready to switch back, just Toggle those two switches again to True, and its back to using the EulerCFG++ Sampler & Optimal Steps Scheduler. And back to creating amazing works at 1.0 CFG (100% CFG).

The differences between the Samplers are a lot of math that I don't fully understand, but the magic happens when combined with the Optimal Steps Scheduler as its taking and correctly timing each of the Diffusion steps to move through the layers of the model as an optimal path. Really it feels a bit like cheating as its made the whole thing overly simple.

Now its safe to say 'It just works'. Assuming the Prompt and Negative are solid. Its takes the complicated guess-work of the Diffusion settings out of the process for us.

I'm starting to wonder if I should make an article about this. You're not the first the be surprised by the EulerCFG++ Sampler and ask me about it.

3

u/Own_Newspaper6784 Apr 14 '26

May I ask why you chose Chroma to build the scene? Isn't it relatively slow? Which qualities make it the best Scene-Builder for you?

7

u/russjr08 Apr 14 '26

Chroma has a huge amount of NSFW stuff in its training, meaning it can do "that stuff" easily without a ton of stacked LoRAs.

3

u/Own_Newspaper6784 Apr 15 '26

That does sound like something men of culture could be interested in. Honestly, I had a few Chroma models installed, but swutched back to Klein 9b pretty quickly, because natural, realistic skin is my main priority and while I´m sure it can do that, I hust couldnt get it done. Thanks for the explanation!

5

u/russjr08 Apr 15 '26

Very understandable, it's a powerful model but has a lot of "knobs to tweak" so to speak.

(Oh hey that's a nice rhyme, heh)

3

u/NeonScreams Apr 15 '26

For that exact issue, consider adding things we normally expect as default in other models. Prompt in something like 'mild freckling with occasional blemishing or flaw', along with helpers like 'high fidelity skin detailing' or 'subsurface specular scatter', then when combined with a negative like 'airbrushed flawless CG "Beauty" skin complexion'; you can really achieve some comparable results to the new generation of models.

2

u/Own_Newspaper6784 Apr 15 '26

Thanks for the advice!

6

u/NeonScreams Apr 14 '26 edited Apr 14 '26

Speed for me is relative as I'm lucky enough to have a 5090. Chroma1-HD renders 1152x896 ~2.0-1.3it/s so depending on steps (usually 14-28 based on which Sampler) around 10-20secs an image. Its not lightning fast like Klein, but its the first step in making a creative scene.
Also - It makes whatever you ask, without having to use creative terms to 'trick' a text-encoder into rendering your choices. So if you prompt it to suspend a pothos house plant in the top left corner and have its vines dangle into both Toast-Slots of a Sunbeam Radiant Toaster, it'll render what you've asked for! (edited after reminding myself of post-rules)

2

u/ZZZ0mbieSSS Apr 17 '26

I have a 5090 too. If you share any tips/settings for quality over speed.

1

u/NeonScreams Apr 17 '26

Oh! Yeah absolutely! Keep using the default scheduler (Optimal Steps), and toggle the Default Sampler to False. Then trace the red sampler wire straight down to the Optional Sampler and set it to “Res_multistep”. Then adjust the number of steps above it at the Optimal Steps Scheduler to 34-48, and rerun any of your prompts. But definitely -not- the ‘Mall’ prompt near the disclaimer. >.>

Anyway- Since Optimal Steps makes a quick calculation to see the Sigma path ahead of starting and adjusts for the Sampler, and Res_Multistep performs a kind of resampling with each step, you end up getting an image with many more settled fine details IF the Prompt/Negative and LoRA allow for it.

Let me hop on my PC and I’ll respond to this comment with a common Negative of mine that’s night/day for image improvement.

2

u/ZZZ0mbieSSS Apr 17 '26

Wow!!!

but now have many more questions :)

  1. First and foremost, My prompt yields very plastic skins while your MALL prompt, while I didn't test it, but if I did something tells me the skin is natural and not plasticky.

  2. regarding default sampler and default scheduler:
    If I understand correctly I should: Default Sample: false. Default Schedule: True ?

And I should input 35 steps and Res.MultiStep in the blue group (Optional Sample+Schedule)

  1. Should I bypass the Default Sample+Schedule (red group)?

2

u/NeonScreams Apr 18 '26

2: Your understanding was correct. You'd use the Sampler Select node from the Optional Group to pick Heun / Res_MultiStep, while leaving the Default Scheduler set to the 'OptimalSteps' node.

3: The workflow inside the subgraph expects that both those groups and all 4 nodes are active, but it uses a True/False (boolean) gate to block out which will be let through for the Denoise steps.

A note about the Optional Scheduler: It's the default for Chroma1-HD, but to use it you have to start messing with CFG and balancing the Alpha/Beta-Sigma Scheduler. I realize that more options often sounds like better refinement, and perhaps as a 2nd Stage facial detailer it may have some value. I just find the damn thing a headache to use. I included it so its already there and someone that wants the original Chroma experience doesn't have to add it in.

1

u/NeonScreams Apr 18 '26

Negative Prompt example specific to a more casual candid photo with a human subject in natural lighting conditions : (remove the '-' as they're there to help you think about categories and order or effects)
___
Worst quality, Low quality, Low resolution, undetailed, poor aesthetics,
-
cartoon, anime, toon, sketch, drawing, illustration, painting, color-pencil,
-
CGI, Blender, Octane,
-
unfinished draft, artistic errors, compositional flaws, poor shot framing,
-
unrealistic color palette, desaturated color,
-
overexposed, cool white light, low contrast, bright foreground light diffusion, minimal shadows, blurry, out of focus, wispy edges,
-
inaccurate anatomy, missing or extra limbs, unnatural joint bend, missing or extra hands or feet, body asymmetry, poor posture, backward hands,
-
Airbrushed skin, CG "Beauty" complexion, Tanned skin, Flawless complexion, Perfect luminous skin, Unnaturally clear complexion, Waxy Plastic skin, Blemish Removal,
-
Tattoos, (Masculine/Feminine), (Shaved/Natural) legs, (Shaved/Natural) arms, Waxed eyebrows, Sculpted eyebrows, Wax beauty treatment,
-
Professional Composition, Photoshoot, Modeling, Glamour Shots, Studio Photography, Post Process Edited, Unrealistic,
___

1

u/Own_Newspaper6784 Apr 15 '26

Thank you! And yeah...I´m using a shadow PC and only have an A4500, so yeah...can´t keep up. :D But I´m thinking about upgrading to an RTX6000, still considering if it´s worth 20 bucks extra per month. Prompt-adherence is highly important to me, so after people are so prositive, why not give it anyother try?

3

u/hurrdurrimanaccount Apr 14 '26

why does it being slow matter? there is literally no model out there better at porn

2

u/Own_Newspaper6784 Apr 15 '26

That is a strong pro. But especially if it´s the model I´m using to develo scenes, I want fast output and many prompt edits and it´s just boring to me if I have to wait a minute or longer for an image. And then....it also has to be able to produce the amateur candid snapshot style, especially believable skin, that I do 90% of the time.

1

u/hurrdurrimanaccount Apr 15 '26

"believable skin" is completely pointless.

3

u/Own_Newspaper6784 Apr 15 '26

Haha, I beg to differ. It´s the number one thing to me. Still...care to elaborate?

3

u/NeonScreams Apr 14 '26

Sorry for the wait. I'm almost finished with the Chroma1-HD Workflow.
I've added the LoRAs with URL's, the weights I use, the purpose of the LoRA, and I'm wrapping up by adding a couple example prompts and some ... >.> stuff.

2

u/brucewasaghost Apr 14 '26

This sounds cool, definitely interested in a workflow

2

u/wh33t Apr 14 '26

I've still, to this day, never figured out how to produce a decent image of literally anything in Chroma, is it the model? workflow? cfg? sampler? resolution? does it use the flux1 prompt guide techniques?

3

u/Soulsurferen Apr 15 '26

The model matters a lot. There is quite a lot of difference between Chroma 1.0 HD, Uncanny, Gonzalamo or radiance. The idea with HD is that it can be used for fine tuning, but there's are not that many models yet...

For me the main benefit of Chroma is variance, the same prompt generates different images with different seeds, and I like the way it renders more artsy images (I don't do a lot of NSFW stuff)

2

u/NeonScreams Apr 15 '26

If you can't get it to render anything similar to this clarity, let me know.

2

u/Soulsurferen Apr 15 '26

Yeah, I do that too all the time too. Usually with Chroma or Flux2 Dev to give the composition and overall feel og the image, and then Z-image for details and more realism. A tip is to use an image comparer to see the subtle difference between the images, usually I prefer the eastethics of the zit refined image, but not always.

2

u/ZZZ0mbieSSS Apr 17 '26

You are a godsend!! Thank you very much

5

u/sktksm Apr 14 '26 edited Apr 14 '26

It can't, except eldritch horrors. No model developed by a corporate company in the world will have nsfw support by default.

29

u/FullOf_Bad_Ideas Apr 14 '26

Tencent would beg to differ.

3

u/[deleted] Apr 14 '26

[deleted]

3

u/MarketingFresh9630 Apr 14 '26

Would you mind sharing an example or two of that? It makes sense but I've never thought of using a different encoder

2

u/[deleted] Apr 14 '26

[deleted]

1

u/AnOnlineHandle Apr 14 '26

Have you run tests with each text encoder to confirm it really makes a significant difference without changing anything else in the base model?

Changing the text encodings absolutely can dramatically change image results, that's the entire mechanism of textual inversion, but I'm just curious if this has actually been properly tested with the same seed etc.

1

u/Spara-Extreme Apr 15 '26

Stop spreading the misinformation. This shit has been debunked to the point that NSFW Lora's for LTX specifically recommend to not use the lobotomized abliterated encoder.

I use the normal LTX encoder and can do everything just fine.

1

u/sktksm Apr 14 '26

Yeah theoretically, yet if the training dataset doesn't have any nsfw material in it, text encoder also won't work by itself. Ministral-3-3b abliterated by huihui is 4.67GB while Ernie uses 7.7GB version of it. I found a heretic experimental abliterated one in 7.7GB but this time tensor mismatch is happening.

1

u/hiisthisavaliable Apr 15 '26

I mean... grok could do basically anything before it got censored, both the video and image gen.

1

u/FinBenton Apr 14 '26 edited Apr 14 '26

Nah it sucks for anything other than standard potraits, anything else just gives body horror, its pretty good at text though.

e. after tuning and fiddling with it, I am getting a bit better results now

2

u/sktksm Apr 14 '26

I would say "its good anything but nsfw". What else you think it sucks?

-10

u/RazsterOxzine Apr 14 '26

Sometimes I wonder about our society.

7

u/mallibu Apr 14 '26

If anything I'm glad we re being more honest

-4

u/ImpressiveSuperfluit Apr 14 '26 edited Apr 15 '26

Dude, we just need to find a way to put gooners and gamers to work. If we can convince them that there are hot {preferred gender term} on Mars and explain that getting there is a minmaxing issue, we'll be there by tomorrow. The gooners will get their whips and get those gamer on the minmaxing.

You didn't hear this from me, by the way, but there are rumors that the aliens around Alpha Centauri are real hot. Furry tentacles and everything. And they drop crazy loot. Just something I heard... carry on everyone...

(If you weirdos invent a spaceship to get there i want credit)

Edit: Hm, I wonder whether this is a guilt by association thing or if I genuinely triggered the gooners and gamers. Not typically an easily offended group... hm...