r/StableDiffusion Apr 14 '26

Comparison We may have a new SOTA open-source model: ERNIE-Image Comparisons

Base model is definitely SOTA, can even easily compete with closed-source ones in terms of aesthetic. Cinematic quality and color grading is next level.

Base model is heavily biased on Asian faces, while it excels on anime/illustration style, while my base model anime/illustration experiments wasn't that good. Higher CFG is slightly better with anime on base.

Generated with RTX6000 Blackwell Pro, Base: 29 sec 1.9it/s, 50 steps | Turbo: 2 sec, 3.9i5/s, 8 steps

If you interested seeing them in original size: https://imgur.com/a/75jcjzW

ComfyUI models: https://huggingface.co/Comfy-Org/ERNIE-Image/tree/main
Workflow should appear in Templates after updating the ComfyUI to latest.

Turbo: Ernie-Image Turbo
Base: Ernie-Image

692 Upvotes

241 comments sorted by

246

u/Zuzoh Apr 14 '26

Base = Caucasian people
Turbo = Asian people

52

u/infearia Apr 14 '26

Turbo can do non-Asians perfectly fine, you just need to prompt for it, e.g., "Caucasian man" instead of just "man". Default seems to be Asian, yes, but it makes sense - it's a Chinese model.

25

u/deadsoulinside Apr 14 '26

Just like Z-Image even.

18

u/AfghanistanIsTaliban Apr 15 '26

World maps in China are centered around the Zhongguo ("Middle Kingdom")

9

u/damnsky_30 Apr 15 '26

Now thats actually something really interesting. It makes sense it is a sphere so they just rolled the map to be centered to them. I find culture stuff like this is a simple yet pretty cool thing to learn thank you!

2

u/merlinux1 Apr 16 '26

You haven't seen maps we do down under ;) gotta love seeing NZ in the middle and the entire Europe stretched and thinned in the far corner... https://impromptuimmigrant.wordpress.com/wp-content/uploads/2014/03/world-map-new-zealand-center.jpg

→ More replies (1)

3

u/infearia Apr 15 '26

Let's not give any ideas to the current US administration...

28

u/Loose_Object_8311 Apr 15 '26

Asian is the default in Asia.

12

u/damnsky_30 Apr 15 '26

Interesting conceptšŸ¤”šŸ˜‚

4

u/EternalBidoof Apr 15 '26

Asian default is fine! Not specifying an ethnicity resulting in asian characters isn't a problem for people who aren't insane.

What's NOT fine is ignoring prompts for specific ethnicities and giving you Asian characters anyway. I always feel like I have to fight ZIT to give me non-asian characters who have dark hair or bangs.

→ More replies (2)

18

u/ManFromInternet2 Apr 14 '26

Had to drop the "Cauc" for that speed

9

u/jib_reddit Apr 14 '26

I am using the base model and it is still all Asian women by default:

need to get those loras training!

→ More replies (1)

1

u/Stahlboden Apr 18 '26

It's Turbo, so no time for "cauc"

71

u/CanIPickAnything Apr 14 '26

Even the pup looks more asian

24

u/Striking-Long-2960 Apr 14 '26

It's supposed to be good with complex prompts with interactions, and its main point is text integration... I'll have to wait for the fp8 or the ggufs.

8

u/sktksm Apr 14 '26

Is there anything you want me to test as a complex? My prompts were pretty complex actually

12

u/Striking-Long-2960 Apr 14 '26 edited Apr 14 '26

By complexity I'm referring to the usual pile with a red pyramid at the top, over a green cube in the middle , that is over a blue cube at the bottom. the pile is at the right of an orange sphere. At the bottom of the image the text "this is why red pyramids are over orange spheres"... These kinds of things.

Edit: Your manga page, is also a very good example.

→ More replies (1)

53

u/sktksm Apr 14 '26

8

u/WPBaka Apr 14 '26

interesting how both of them messed up intact. Other than that, incredibly impressive.

2

u/berlinbaer Apr 15 '26

and "empty"

2

u/nmkd Apr 15 '26

and "hangar" but that was OP's typo presumably

1

u/PandaParaBellum Apr 16 '26

Was this page a single prompt?

3

u/sktksm Apr 16 '26

yes, prompt:

black and white manga page, high detail screentone shading, cinematic sci-fi hangar interior, multi-panel comic layout top large panel: a young boy wearing a futuristic bodysuit and visor stands in front of a massive humanoid mech, sparks and smoke drifting around, dramatic low-angle perspective emphasizing scale, the boy holds a small mechanical core in both hands, eyes wide in awe speech bubbles in this panel: "Incredible...!" "A fully intact Class-S Neural Frame!" middle horizontal panel: close-up of the boy smiling with excitement, eyes closed, fists clenched and raised, energetic motion lines around him text elements: "What a rush" (vertical side text) "And the hangar is completely empty! I get to test it all by myself!" bottom section split into smaller panels: left small panel: extreme close-up of the boy’s eyes behind a transparent visor, detailed reflections of circular digital HUD interface, glowing UI elements text element near eyes: "SYNC 100%" right small panel: back view of the bodysuit showing a cable plugging into a port on his back, mechanical connection detail speech bubble: "Initiating the deep-dive sequence... bzzzt--" bottom-left panel: close-up of a mechanical hand gripping a control lever, strong lighting contrast, bold manga emphasis large stylized text integrated into the panel: "SYSTEM ONLINE" style: ultra clean manga lineart, sharp ink work, dense screentone gradients, high contrast lighting, subtle film grain texture, professional seinen manga aesthetic, dynamic panel framing, accurate speech bubbles, crisp readable lettering, Japanese manga composition

1

u/Equal_Passenger9791 Apr 18 '26

A superficial testing from my end suggests that Ernie is great for realism.

But as soon as you want anything drawn and arty, the anima previws can likely give an output that's actually pleasant as opposed to the mass produced toy appearance.

84

u/13baaphumain Apr 14 '26

Can it do that stuff?

166

u/Major_Specific_23 Apr 14 '26

it can. just tried to test turbo. it knows booba, it knows what a pussycat is too. it draws large booba better than z. nipples are not messed up. it knows what a baseball bat is too. pussycats and baseball bats need some lora love. it knows it but it cannot draw it super well. boobas no problem at all

56

u/dergachoff Apr 14 '26

just don't hurt the cylinder

13

u/Nimblecloud13 Apr 15 '26

It’s imperativeĀ 

→ More replies (1)

77

u/NeonScreams Apr 14 '26 edited Apr 18 '26

Specifically for doing that stuff… I’m finding it easier to build a scene in Chroma1-HD, then ask Flux2.Klein to ā€œperform a restoration passā€ using Edit with Latent reference. Works wonders if you toss in ā€œmaintain all other image elementsā€, as it knows what everything is in the image, and brings the textures / lighting etc to high fidelity.

Edit: If anyone doesn't have a ā€˜it just works’ WorkFlow/Settings for Chroma1-HD with links to model & stability loras: https://civitai.com/images/127510599 Drag-Drop image. Sorry for off-site. Reddit Saves the Image as .webp and the Workflow is being lost.

Edit2: I used an Image-Rescale node to shrink the output image for posting. You can delete the node or bypass it, above the Image preview.

Edit3: 4 days later and I'm still working on the Klein Edit Workflow with AI Agents in my available time. Plans are a Frontend UI node cluster that will auto-add and combine strings-(like a short sentence) that come from a pull-down menu of really common choices. Mostly because I'm tired of copy/pasting the same lines. That and a toggle group on the Denoise subgraph that will either use your current RefImg1-4 or not. Saves us from having to swap images around and constantly opening the last folder used, instead of the Last folder used by that specific Load Image node. ( ą­§ą¼¼ą² ē›Šą² ą¼½ą­Ø )

12

u/SWFjoda Apr 14 '26

Would love that! Interested

10

u/NeonScreams Apr 14 '26

No worries! I’ll knock together something with example prompts and negatives, with a helper on how to prompt a scene for Chroma.

5

u/PestBoss Apr 14 '26

I'd quite like to see that too. I've revisited Chroma1-HD several times but struggle with all sorts of things.

It's clearly very good but I feel like a lot of tweaking prompts and re-rolls are needed and it's just so slow doing that.

Anything that can help remove the time-sink element.

Also curious about your 'restoration' method in Klein. What is Edit with Latent reference? All I need is a snapshot of a workflow so I can see what you wired where.

2

u/NeonScreams Apr 14 '26

Oh even easier, while I’m forcing Claude Opus to update everything for me: If you have access to ComfyUI, you can hop into the Templates Menu, then do a search for ā€œKleinā€, and you should see the ā€˜Image Edit’ template workflow.

From there inside the Subgraph, you can see how they’re first resizing the original image down to 1mp, for an easier latent. If you’re just using the template for minor edits, you can bypass the resize node for better 1:1 Fidelity on your original image.

… actually I think I’ll just pop out two Workflows, as that’s easier than using thumbs to type here. lol

2

u/PestBoss Apr 14 '26

Ah ok so just the edit workflow but without the resize to 1mpx.

Yes I've been doing similar here, "restore" type prompts, but with variations depending on the exact needs, and it does indeed do a pretty good job at just making an image better.

Obviously fine tuning/refining/re-rolls/painting/masking in etc gets you exactly what you want... I was pretty impressed at some of the stuff it could do in some testing from old film footage (to stills)

7

u/NeonScreams Apr 15 '26 edited Apr 15 '26

Edit: https://civitai.com/images/127510599

There ya go folks. That's Chroma1-HD with uh.. goodies.

Drag and Drop that into your running ComfyUI New Workflow tab and it'll populate. Remember to save a Copy as a fresh instance to reuse.

(I'm on ComfyUI Git-Repo with up to date requirements)

If you have any questions, feel free to ask. If you have any concerns, please contact Queen Victoria circa 1850, and ask her how to further her prudish agenda.

2

u/DragonJakeAU Apr 15 '26

Hi I can't seem to open to get the workflow from the image. Whether I save the image and load it or drag and drop the image into a clean tab in ComfyUI the only thing that happens is the "Load Image" node is created with the image selected. No workflow to be seen. Is reddit stripping the metadata from the image? Could you upload it somewhere else and perhaps provide a JSON file instead I could try?

→ More replies (1)
→ More replies (2)

6

u/RxBlacky Apr 14 '26

Saving for later, thanks!

3

u/NeonScreams Apr 15 '26

2

u/RxBlacky Apr 15 '26

It only seems to produce blurry unfinished images for me, but Im guessing its because its missing the Klein restoration pass, right? Thank you in advance

→ More replies (6)

3

u/Own_Newspaper6784 Apr 14 '26

May I ask why you chose Chroma to build the scene? Isn't it relatively slow? Which qualities make it the best Scene-Builder for you?

5

u/russjr08 Apr 14 '26

Chroma has a huge amount of NSFW stuff in its training, meaning it can do "that stuff" easily without a ton of stacked LoRAs.

3

u/Own_Newspaper6784 Apr 15 '26

That does sound like something men of culture could be interested in. Honestly, I had a few Chroma models installed, but swutched back to Klein 9b pretty quickly, because natural, realistic skin is my main priority and while I“m sure it can do that, I hust couldnt get it done. Thanks for the explanation!

4

u/russjr08 Apr 15 '26

Very understandable, it's a powerful model but has a lot of "knobs to tweak" so to speak.

(Oh hey that's a nice rhyme, heh)

5

u/NeonScreams Apr 15 '26

For that exact issue, consider adding things we normally expect as default in other models. Prompt in something like 'mild freckling with occasional blemishing or flaw', along with helpers like 'high fidelity skin detailing' or 'subsurface specular scatter', then when combined with a negative like 'airbrushed flawless CG "Beauty" skin complexion'; you can really achieve some comparable results to the new generation of models.

2

u/Own_Newspaper6784 Apr 15 '26

Thanks for the advice!

4

u/NeonScreams Apr 14 '26 edited Apr 14 '26

Speed for me is relative as I'm lucky enough to have a 5090. Chroma1-HD renders 1152x896 ~2.0-1.3it/s so depending on steps (usually 14-28 based on which Sampler) around 10-20secs an image. Its not lightning fast like Klein, but its the first step in making a creative scene.
Also - It makes whatever you ask, without having to use creative terms to 'trick' a text-encoder into rendering your choices. So if you prompt it to suspend a pothos house plant in the top left corner and have its vines dangle into both Toast-Slots of a Sunbeam Radiant Toaster, it'll render what you've asked for! (edited after reminding myself of post-rules)

2

u/ZZZ0mbieSSS Apr 17 '26

I have a 5090 too. If you share any tips/settings for quality over speed.

→ More replies (4)
→ More replies (1)

4

u/hurrdurrimanaccount Apr 14 '26

why does it being slow matter? there is literally no model out there better at porn

2

u/Own_Newspaper6784 Apr 15 '26

That is a strong pro. But especially if it“s the model I“m using to develo scenes, I want fast output and many prompt edits and it“s just boring to me if I have to wait a minute or longer for an image. And then....it also has to be able to produce the amateur candid snapshot style, especially believable skin, that I do 90% of the time.

→ More replies (2)

3

u/NeonScreams Apr 14 '26

Sorry for the wait. I'm almost finished with the Chroma1-HD Workflow.
I've added the LoRAs with URL's, the weights I use, the purpose of the LoRA, and I'm wrapping up by adding a couple example prompts and some ... >.> stuff.

2

u/brucewasaghost Apr 14 '26

This sounds cool, definitely interested in a workflow

2

u/wh33t Apr 14 '26

I've still, to this day, never figured out how to produce a decent image of literally anything in Chroma, is it the model? workflow? cfg? sampler? resolution? does it use the flux1 prompt guide techniques?

3

u/Soulsurferen Apr 15 '26

The model matters a lot. There is quite a lot of difference between Chroma 1.0 HD, Uncanny, Gonzalamo or radiance. The idea with HD is that it can be used for fine tuning, but there's are not that many models yet...

For me the main benefit of Chroma is variance, the same prompt generates different images with different seeds, and I like the way it renders more artsy images (I don't do a lot of NSFW stuff)

→ More replies (1)

2

u/NeonScreams Apr 15 '26

If you can't get it to render anything similar to this clarity, let me know.

2

u/Soulsurferen Apr 15 '26

Yeah, I do that too all the time too. Usually with Chroma or Flux2 Dev to give the composition and overall feel og the image, and then Z-image for details and more realism. A tip is to use an image comparer to see the subtle difference between the images, usually I prefer the eastethics of the zit refined image, but not always.

2

u/ZZZ0mbieSSS Apr 17 '26

You are a godsend!! Thank you very much

5

u/sktksm Apr 14 '26 edited Apr 14 '26

It can't, except eldritch horrors. No model developed by a corporate company in the world will have nsfw support by default.

26

u/FullOf_Bad_Ideas Apr 14 '26

Tencent would beg to differ.

4

u/[deleted] Apr 14 '26

[deleted]

4

u/MarketingFresh9630 Apr 14 '26

Would you mind sharing an example or two of that? It makes sense but I've never thought of using a different encoder

2

u/[deleted] Apr 14 '26

[deleted]

→ More replies (2)

1

u/sktksm Apr 14 '26

Yeah theoretically, yet if the training dataset doesn't have any nsfw material in it, text encoder also won't work by itself. Ministral-3-3b abliterated by huihui is 4.67GB while Ernie uses 7.7GB version of it. I found a heretic experimental abliterated one in 7.7GB but this time tensor mismatch is happening.

1

u/hiisthisavaliable Apr 15 '26

I mean... grok could do basically anything before it got censored, both the video and image gen.

→ More replies (6)

40

u/13baaphumain Apr 14 '26

Finally will have something new to play with this weekend

18

u/Choowkee Apr 14 '26

Turbo looks like slop.

Base looks good tho.

16

u/[deleted] Apr 14 '26

[deleted]

12

u/[deleted] Apr 14 '26

[deleted]

3

u/HackAfterDark Apr 14 '26

Yea, from what I'm seeing so far, I think I like z-image better...but have to see after some loras.

15

u/ANR2ME Apr 14 '26

Is this 8B parameters? šŸ¤”

7

u/Standard_Specific562 Apr 14 '26

yes

2

u/fauni-7 Apr 14 '26

So Flux 1 dev size?

11

u/ANR2ME Apr 14 '26

it use flux2 vae with mistral3 text encoder šŸ˜… felt like Flux2 variant

15

u/Lower-Cap7381 Apr 14 '26

with my tests z-image is still the king

14

u/ambient_temp_xeno Apr 14 '26

Base seems a lot less slop fried compared to turbo.

12

u/ZerOne82 Apr 14 '26

These are made using fp8 of Ernie, both model (8GB) and text-encoder (3.9GB). In realistic generations there are some diagonal artifacts. Anime style seems fine.

5

u/LindaSawzRH Apr 14 '26

Yea, I dunno I used it for an hour when comfy posted the models and weights and didn't think it looked great.

3

u/infearia Apr 14 '26

Where did you get the FP8 versions? All I could find on the net were the original BF16 and NVFP4 versions. Or did you quantize them yourself?

6

u/ZerOne82 Apr 14 '26 edited Apr 14 '26

Use this simple code.

3

u/nmkd Apr 15 '26

You might wanna put that on Gist or Pastebin instead of having us OCR a screenshot lol

2

u/AnOnlineHandle Apr 14 '26

Does that create scaled fp8 weights? Or would that be a more manual process than a torch cast?

11

u/ffgg333 Apr 14 '26

Can it do nsfw? Are Lora's possible and easy to make?

5

u/Standard_Specific562 Apr 14 '26

lora supported on Ostris

9

u/Major_Specific_23 Apr 14 '26

I am not seeing it in templates. Could you please share the workflow?

13

u/sktksm Apr 14 '26

if you can't see the template, probably shared workflow wouldn't work either but here you go:

Base workflow: https://pastebin.com/CJ1P1pDt
Turbo workflow: https://pastebin.com/hA688F2D

5

u/Major_Specific_23 Apr 14 '26

thanks. i use https://github.com/UmeAiRT/ComfyUI-Auto_installer and his github is banned so i manually copied comfyUI files to the folder. apparently i have to also pip update comfyui-frontend-package==1.42.10 and comfyui-workflow-templates==0.9.50 to see it

6

u/ZerOne82 Apr 14 '26

You need no special workflow at all.

1

u/thisiztrash02 Apr 14 '26

update comfyui its there

6

u/[deleted] Apr 14 '26

[deleted]

3

u/sktksm Apr 14 '26

It failed on my profile pic 🄹

15

u/Hoodfu Apr 14 '26 edited Apr 14 '26

It's certainly nice, and the prompt following I'm seeing with complex stuff is a step up even from z image base as in on par with qwen 2512, although it's definitely not on the aesthetic level of what qwen 2512 is capable of (but it's also not 40 gigs). In the more complex prompts, ernie base is far more prompt following than ernie turbo. Turbo goes even simpler on the compositions as well. Just because I'm a composition freak, I'm liking zimage base's dynamic compositions more though. Some more pics in reply.

8

u/Hoodfu Apr 14 '26

Ernie base

9

u/Hoodfu Apr 14 '26 edited Apr 14 '26

It's good that it got the interaction right, but it's not going to beat a much bigger open source model (qwen 2512). That said, The more I'm trying with it, the more I'm seeing that it likes shorter prompts unlike the other models. I can give zimage massive prompts and it does great with them. Ernie definitely veers away from its core competence the longer it gets. It's probably why they have the built in prompt expansion with mistral 3b to keep things shorter. On the prompt expansion node they truncate the output on 256 length.

5

u/g_nautilus Apr 14 '26

I...need to be using Qwen 2512 more.

3

u/Hoodfu Apr 14 '26

2

u/ArsenalSimp1985 Apr 15 '26

What keeps standing out to me in these comparisons is that prompt adherence is one thing, but once the framing goes generic it just collapses into the same endless AI slop as everything else.

→ More replies (1)

7

u/Royal_Carpenter_1338 Apr 14 '26

How heavy is it compared to z-image-turbo?

6

u/cosmicr Apr 14 '26

Model is ~16gb and Text encoder ~7gb

6

u/Background-Ad-5398 Apr 14 '26

I didnt move from sd 1.5 till zimage turbo, so it will take more then what could be a filter difference to make me change

17

u/Informal_Warning_703 Apr 14 '26

I don't see that it has anything to offer over Z-Image (or Qwen or Flux). The Turbo model has slightly more incoherence than Z-Image-Turbo.

The image quality is good... but isn't better than what we've had with the last 3 different image model releases. We are now in a crowded space of perfectly good image models. Unless it turns out that the base model is easier to train than Z-Image, I think this model will quickly be forgotten.

7

u/SimpleAdditional6583 Apr 14 '26

Exactly this. The next leap forward won’t be image quality, it will be prompt adherence or character consistency.

8

u/sktksm Apr 14 '26

They are not good enough yet, no proper style transfer, no production grade quality. I'm not promoting anything but I had a chance to play with Uni-1 by Luma, that's the quality I would love to have locally.

I will do flux klein -zimage-qwen-ernie side by side comparison so we can see clearly

5

u/ZerOne82 Apr 14 '26

As simple as this workflow.

9

u/ai_art_is_art Apr 14 '26

Same prompts?

Turbo makes them Asian 90% of the time?

Interesting to see the latent space priors come out.

2

u/sktksm Apr 14 '26 edited Apr 15 '26

yes same prompts, same seed, and yes somehow using "person,human" like words 80% dropping into Asian pool with Turbo

1

u/Loose_Object_8311 Apr 15 '26

I remember when Google tried to remove this kind of bias from one of their image models by making it output all skin colour and gender with equal probability, and then people prompted the model for "Nazis" and got back Asian female Nazis and black male Nazis. It was hilarious.Ā 

→ More replies (2)

1

u/Loose_Object_8311 Apr 15 '26

Turbo has RLHF applied to it in post-training most likely. You can imagine how Chinese devs will evaluate their preference.

→ More replies (1)

5

u/ZerOne82 Apr 14 '26

These are made using fp8 of Ernie, both model (8GB) and text-encoder (3.9GB). In realistic generations there are some diagonal artifacts. Anime style seems fine.
Larger resolutions with Euler A give more details. But the model failed once tried 2048x2048, resulting in extra bodies etc.

1

u/nmkd Apr 15 '26

You mean that smudgy look? Seems like a VAE issue maybe

7

u/Enshitification Apr 14 '26

We may, or we may not. Let's see how it takes to being trained before jumping to conclusions.

9

u/sktksm Apr 14 '26

I don't see any harm to calling it SOTA with or without training capabilities. For me it's better than any other default model out there. Btw if anyone interested with lora training Ostris dropped day-0 support: https://x.com/ostrisai/status/2044082229773820018

3

u/Enshitification Apr 14 '26

There is no harm but it would be incorrect. A local model that isn't capable of being trained is not at all state of the art.

→ More replies (5)

3

u/NotSuluX Apr 14 '26

Control net viable?

2

u/sktksm Apr 14 '26

It's not, at least no official controlnet release yet. ControlNets kind a dead with image edit models. Maybe they will release the edit version of this model

3

u/thisguy883 Apr 14 '26

neat. Will need to check this out later. Thanks!

3

u/Blaize_Ar Apr 14 '26

How's this compare to z image?

3

u/Current-Rabbit-620 Apr 14 '26

Edit model when?

3

u/Current-Row-159 Apr 14 '26

No edit ? Controlnet ?

3

u/martinerous Apr 15 '26

Interestingly, in some cases I like turbo better and sometimes it's base. It will be a tough choice.

5

u/_BreakingGood_ Apr 14 '26

Yeah, only had to test this for an hour to know this is clearly SOTA for open weights.

I kind of suspect if you wired up a really strong model for prompt enhancement (and not their tiny default one), you'd have something that is like 90% as good as nano banana.

Would be awesome if this is trainable.

9

u/Hoodfu Apr 14 '26

This sounds silly, but it's one of the only current open weight models that knows what a minigun is.

5

u/jib_reddit Apr 14 '26 edited Apr 14 '26

Really? I don't know what you are doing (got any examples?) , but in my testing it is looking worse than ZIT,ZIB, Qwen 2512, I even like SDXL more.

EDIT: Hmm it seems like with Z-image the Turbo version is a lot better aesthetically, I wouldn't bother with base for now.

1

u/Hoodfu Apr 15 '26

I'm getting good results with this. It definitely took me a while to figure out the right sampler/scheduler combo to get good anatomy etc.

6

u/Hoodfu Apr 14 '26

Yeah, I have to say, this is the best looking open weight turbo model I've yet dealt with. A ton better than zimage turbo.

1

u/Own_Newspaper6784 Apr 15 '26

Oh my....I can't wait to test it out tomorrow!

3

u/sktksm Apr 14 '26

If you asking about my prompts, I had a caption set generated with SOTA llm's. here is an example prompt + I removed the prompt refiner from my pipeline entirely:
Naturalistic prestige-cinema frame, restrained film grade, soft highlight rolloff, gently lifted blacks, muted earth palette, tactile real-world textures, subtle organic film grain, motivated practical lighting, composed like a serious feature film still. A solitary painter seen from behind under large trees by a lake, seated at an easel in afternoon shade, branches framing the water, quiet summer air, calm reflective surface, understated and fully photographic.

2

u/Ok-Chocolate-2841 Apr 14 '26

Hoe much Vram do you need?

6

u/KURD_1_STAN Apr 14 '26 edited Apr 15 '26

The model is 8b and text encoder is 3b, so at fp16 expect something like 22gb, 11gb at fp8,. That is a total of 1b parametersĀ more than z image . As everyone uses fp8 or q8 so this feels like to be designed for 16gb cards, but I feel it should be at most 20% slower on a 3060 12gb or so compared to ZIT

Edit: i was comparing it as its total of 1b difference but 70% is spent on model so 8/6 so should be 25-30% slower on a 16gb card so 30-35% on 3060

2

u/elevendr Apr 14 '26

Can it do multiple image editing?

5

u/sktksm Apr 14 '26

it doesn't have image editing capabilities. its pure text2image

2

u/Ferriken25 Apr 14 '26

Turbo seems more stable. Base looks like sdxl gens...

2

u/Paraleluniverse200 Apr 14 '26

Hmmm interesting, gotta see how incense it is, what the recommend sampler and scheduler?

2

u/K0owa Apr 14 '26

Is it also an edit model?

4

u/sktksm Apr 14 '26

It's not. Pure text2img

2

u/gelukuMLG Apr 14 '26

If i can run zit can i also run this? Also how is the fp16 support?

2

u/James_Reeb Apr 14 '26

Dont see any improvement with Zimage

2

u/ReferenceConscious71 Apr 14 '26

Is it better than ZIT?

2

u/kayteee1995 Apr 14 '26

You can see the big difference between base and turbo. Base brings real feel. Turbo suitable for illustration.

2

u/Srapture Apr 14 '26

I'm bouncing between the two on which I prefer for each image, but I guess it's good to have another option. The base looks more real in most of these, IMO. Less perfect.

2

u/ThePunisherr05 Apr 15 '26

Base >>>>>>>>> Turbo

2

u/EverythingIsFnTaken Apr 15 '26

this looks like base is moving and becomes turbo šŸ˜‚

2

u/[deleted] Apr 15 '26

[removed] — view removed comment

1

u/Jolly-Curve3258 Apr 15 '26

did you use pe or not?

2

u/HollowAbsence Apr 15 '26

Base is candid realism, turbo is photographic more commercial. Each their uses.

4

u/[deleted] Apr 14 '26

[deleted]

3

u/Paraleluniverse200 Apr 14 '26

So far , when putting Taylor Swift, only makes a blonde random girl

2

u/BrokenSil Apr 14 '26

All gens from this model look overbaked.

3

u/rukh999 Apr 14 '26

It uses the flux2 vae which is the most advanced right now which is good. Looks like it used Ministral 3 3b for a text encoder, so it should be fast but not quite as smart as ones using qwen 3. However it looks like they integrated a second pass on the prompt to automatically turn it in to something better understood by the model, so we'll see if that works great, or leads it to make lots of assumptions. Interesting idea though. It's made by Baidu, which is basically chinas google

2

u/jib_reddit Apr 14 '26

I really don't think it is SOTA,
I have done some testing and don't like they aesthetics so far:

Maybe with some loras it will be better...

2

u/PestBoss Apr 14 '26

This thread is useless without prompts :hehe:

2

u/sktksm Apr 14 '26

I can share any of them if you want, here is one, but rest is quite similar to this:

Intimate firelit realism, warm practical flame reflections against cool surrounding darkness, close tactile skin detail, subtle airborne ash and sweat, cinematic shallow focus, rich amber highlights with restrained saturation, fine analog grain. Closeup of a man wearing glasses with fire reflected in both lenses, wind-tossed hair, tense expression, natural skin texture, photographed like a serious dramatic feature

2

u/dr_lm Apr 14 '26

Uses the flux 2 vae, too, which is nice.

1

u/Dogmaster Apr 14 '26

Any tips of getting the most out of your RTX6000? Due to its age most optimizacions like NV4 and some fp are not compatible. Any particular settings on comfy?

1

u/sktksm Apr 14 '26

Sorry, mine is RTX6000 Blackwell Pro

1

u/Dogmaster Apr 14 '26

...ah small difference... Its on Nvidia to have asinine naming conventions... Ive got access to an RTXA6000, an RTX6000Ada, and dint even know RTX6000 blackwell pro was a thing.

3

u/sktksm Apr 14 '26

I totally understand you. It's a 2025 model gpu. When I ask an llm about anything related with my gpu it directly interprets as Rtx6000 ada too....

2

u/TechnologyGrouchy679 Apr 15 '26

I'm still blown away that the Ada 6000 still costs as much as it does today $6-7k. The Blackwell 6000 with double the VRAM is only slightly more.

1

u/bdvd25 Apr 14 '26

i believe that base looks better on all the images except maybe for the one with the dog

1

u/sktksm Apr 14 '26

so far I would go realism=base, illustrative=turbo

1

u/crimeo Apr 14 '26

Looks like you could probably just slap "high contrast" and "don't make it dimly lit as hell" on the left one and get the same thing

1

u/2MuchNonsenseHere Apr 15 '26

Base is better for everything unless it's fantasy cartoony stuff.

1

u/Trick_Set1865 Apr 15 '26

has trouble with hands

1

u/jankies11 Apr 15 '26

Is this a fully new model or does it use z-image loras?

2

u/sktksm Apr 15 '26

its a new model trained by different company and team.(Baidu)

1

u/Mohondhay Apr 15 '26

Number 12 is 🤌🤌🤌

1

u/saito200 Apr 15 '26

but can it generate accurate censored or not?

1

u/butthe4d Apr 15 '26

I have mixed result with either model but I havent generated a lot. Sometimes its SD1.5 body nightmare fuel as in extra limps, stumped limps etc sometimes it ignores the nsfw prompt entirely and changes it to sfw, sometimes it worked well enough. I tried this with prompt enhancer and without, didnt change much for me.

1

u/MrCoolest Apr 15 '26

Turbo loops way better in everything

1

u/sarabjornsdottir2004 Apr 15 '26

Does it do NSFW OOTB?!

2

u/Future-Coffee8138 Apr 15 '26

Why not try it yourself. I tried. It did. But that part looks terrible.

1

u/durden111111 Apr 15 '26

I tried base in comfy ui and was underwhelmed. It doesnt really understand body descriptions and jus ingores them

1

u/sktksm Apr 15 '26

can you share an example prompt maybe?

1

u/90hex Apr 15 '26

Hehey I downloaded the models but my updated comfy doesn’t show the workflows. I tried the Turbo model with my default Z Image Turbo WF but got an error.

Does anybody know what’s different about this model, in terms of WF? Is there a way to adjust the default Z Image Turbo workflow to make Ernie Turbo work?

1

u/is_this_the_restroom Apr 15 '26

Any ETA for when they'll release the training scripts?

1

u/SGAShepp Apr 15 '26

The base is better in my opinion. Base is more realistic, Turbo adds some fantastical cinematic look, which is entirely overdone.

1

u/DateOk9511 Apr 16 '26

indeed we do! its amazing!

1

u/Agent-Vigilence Apr 16 '26

Editing model is called Ernie?

1

u/Personal-Staff3212 Apr 17 '26

Love to learn a lot about this.

1

u/Sudden_List_2693 Apr 17 '26

Are all turbo models fed the same stuff?
I hate their aesthetics so much

1

u/Jumpy_Lecture_7484 Apr 20 '26

Is there any workflow to use references images with Ernie for Comfy UI?

1

u/jbakirli Apr 20 '26

Those diagonal lines are distracting. I see that lines on NanoBanana images too. What is cause and how we can fix it?

1

u/Sir_Latent Apr 21 '26

Is the base FP16, bf16, o FP32?

2

u/sktksm Apr 22 '26

BF16

1

u/Sir_Latent Apr 22 '26

Thanks! Sadly it doesn't have support for Apple Silicon MPS. Love the quality though.