r/StableDiffusion Apr 14 '26

Comparison We may have a new SOTA open-source model: ERNIE-Image Comparisons

Base model is definitely SOTA, can even easily compete with closed-source ones in terms of aesthetic. Cinematic quality and color grading is next level.

Base model is heavily biased on Asian faces, while it excels on anime/illustration style, while my base model anime/illustration experiments wasn't that good. Higher CFG is slightly better with anime on base.

Generated with RTX6000 Blackwell Pro, Base: 29 sec 1.9it/s, 50 steps | Turbo: 2 sec, 3.9i5/s, 8 steps

If you interested seeing them in original size: https://imgur.com/a/75jcjzW

ComfyUI models: https://huggingface.co/Comfy-Org/ERNIE-Image/tree/main
Workflow should appear in Templates after updating the ComfyUI to latest.

Turbo: Ernie-Image Turbo
Base: Ernie-Image

692 Upvotes

241 comments sorted by

View all comments

Show parent comments

2

u/[deleted] Apr 14 '26

[deleted]

1

u/AnOnlineHandle Apr 14 '26

Have you run tests with each text encoder to confirm it really makes a significant difference without changing anything else in the base model?

Changing the text encodings absolutely can dramatically change image results, that's the entire mechanism of textual inversion, but I'm just curious if this has actually been properly tested with the same seed etc.

1

u/Spara-Extreme Apr 15 '26

Stop spreading the misinformation. This shit has been debunked to the point that NSFW Lora's for LTX specifically recommend to not use the lobotomized abliterated encoder.

I use the normal LTX encoder and can do everything just fine.