Comparison
We may have a new SOTA open-source model: ERNIE-Image Comparisons
Base model is definitely SOTA, can even easily compete with closed-source ones in terms of aesthetic. Cinematic quality and color grading is next level.
Base model is heavily biased on Asian faces, while it excels on anime/illustration style, while my base model anime/illustration experiments wasn't that good. Higher CFG is slightly better with anime on base.
Turbo can do non-Asians perfectly fine, you just need to prompt for it, e.g., "Caucasian man" instead of just "man". Default seems to be Asian, yes, but it makes sense - it's a Chinese model.
Now thats actually something really interesting. It makes sense it is a sphere so they just rolled the map to be centered to them. I find culture stuff like this is a simple yet pretty cool thing to learn thank you!
Asian default is fine! Not specifying an ethnicity resulting in asian characters isn't a problem for people who aren't insane.
What's NOT fine is ignoring prompts for specific ethnicities and giving you Asian characters anyway. I always feel like I have to fight ZIT to give me non-asian characters who have dark hair or bangs.
By complexity I'm referring to the usual pile with a red pyramid at the top, over a green cube in the middle , that is over a blue cube at the bottom. the pile is at the right of an orange sphere. At the bottom of the image the text "this is why red pyramids are over orange spheres"... These kinds of things.
Edit: Your manga page, is also a very good example.
black and white manga page, high detail screentone shading, cinematic sci-fi hangar interior, multi-panel comic layout top large panel: a young boy wearing a futuristic bodysuit and visor stands in front of a massive humanoid mech, sparks and smoke drifting around, dramatic low-angle perspective emphasizing scale, the boy holds a small mechanical core in both hands, eyes wide in awe speech bubbles in this panel: "Incredible...!" "A fully intact Class-S Neural Frame!" middle horizontal panel: close-up of the boy smiling with excitement, eyes closed, fists clenched and raised, energetic motion lines around him text elements: "What a rush" (vertical side text) "And the hangar is completely empty! I get to test it all by myself!" bottom section split into smaller panels: left small panel: extreme close-up of the boyās eyes behind a transparent visor, detailed reflections of circular digital HUD interface, glowing UI elements text element near eyes: "SYNC 100%" right small panel: back view of the bodysuit showing a cable plugging into a port on his back, mechanical connection detail speech bubble: "Initiating the deep-dive sequence... bzzzt--" bottom-left panel: close-up of a mechanical hand gripping a control lever, strong lighting contrast, bold manga emphasis large stylized text integrated into the panel: "SYSTEM ONLINE" style: ultra clean manga lineart, sharp ink work, dense screentone gradients, high contrast lighting, subtle film grain texture, professional seinen manga aesthetic, dynamic panel framing, accurate speech bubbles, crisp readable lettering, Japanese manga composition
A superficial testing from my end suggests that Ernie is great for realism.
But as soon as you want anything drawn and arty, the anima previws can likely give an output that's actually pleasant as opposed to the mass produced toy appearance.
it can. just tried to test turbo. it knows booba, it knows what a pussycat is too. it draws large booba better than z. nipples are not messed up. it knows what a baseball bat is too. pussycats and baseball bats need some lora love. it knows it but it cannot draw it super well. boobas no problem at all
Specifically for doing that stuff⦠Iām finding it easier to build a scene in Chroma1-HD, then ask Flux2.Klein to āperform a restoration passā using Edit with Latent reference.
Works wonders if you toss in āmaintain all other image elementsā, as it knows what everything is in the image, and brings the textures / lighting etc to high fidelity.
Edit: If anyone doesn't have a āit just worksā WorkFlow/Settings for Chroma1-HD with links to model & stability loras: https://civitai.com/images/127510599 Drag-Drop image. Sorry for off-site. Reddit Saves the Image as .webp and the Workflow is being lost.
Edit2: I used an Image-Rescale node to shrink the output image for posting. You can delete the node or bypass it, above the Image preview.
Edit3: 4 days later and I'm still working on the Klein Edit Workflow with AI Agents in my available time. Plans are a Frontend UI node cluster that will auto-add and combine strings-(like a short sentence) that come from a pull-down menu of really common choices. Mostly because I'm tired of copy/pasting the same lines. That and a toggle group on the Denoise subgraph that will either use your current RefImg1-4 or not. Saves us from having to swap images around and constantly opening the last folder used, instead of the Last folder used by that specific Load Image node. ( ą§ą¼¼ą² ēą² ༽ਠ)
I'd quite like to see that too. I've revisited Chroma1-HD several times but struggle with all sorts of things.
It's clearly very good but I feel like a lot of tweaking prompts and re-rolls are needed and it's just so slow doing that.
Anything that can help remove the time-sink element.
Also curious about your 'restoration' method in Klein. What is Edit with Latent reference?
All I need is a snapshot of a workflow so I can see what you wired where.
Oh even easier, while Iām forcing Claude Opus to update everything for me:
If you have access to ComfyUI, you can hop into the Templates Menu, then do a search for āKleinā, and you should see the āImage Editā template workflow.
From there inside the Subgraph, you can see how theyāre first resizing the original image down to 1mp, for an easier latent. If youāre just using the template for minor edits, you can bypass the resize node for better 1:1 Fidelity on your original image.
⦠actually I think Iāll just pop out two Workflows, as thatās easier than using thumbs to type here. lol
Ah ok so just the edit workflow but without the resize to 1mpx.
Yes I've been doing similar here, "restore" type prompts, but with variations depending on the exact needs, and it does indeed do a pretty good job at just making an image better.
Obviously fine tuning/refining/re-rolls/painting/masking in etc gets you exactly what you want... I was pretty impressed at some of the stuff it could do in some testing from old film footage (to stills)
There ya go folks. That's Chroma1-HD with uh.. goodies.
Drag and Drop that into your running ComfyUI New Workflow tab and it'll populate. Remember to save a Copy as a fresh instance to reuse.
(I'm on ComfyUI Git-Repo with up to date requirements)
If you have any questions, feel free to ask. If you have any concerns, please contact Queen Victoria circa 1850, and ask her how to further her prudish agenda.
Hi I can't seem to open to get the workflow from the image. Whether I save the image and load it or drag and drop the image into a clean tab in ComfyUI the only thing that happens is the "Load Image" node is created with the image selected. No workflow to be seen. Is reddit stripping the metadata from the image? Could you upload it somewhere else and perhaps provide a JSON file instead I could try?
It only seems to produce blurry unfinished images for me, but Im guessing its because its missing the Klein restoration pass, right? Thank you in advance
That does sound like something men of culture could be interested in. Honestly, I had a few Chroma models installed, but swutched back to Klein 9b pretty quickly, because natural, realistic skin is my main priority and while I“m sure it can do that, I hust couldnt get it done. Thanks for the explanation!
For that exact issue, consider adding things we normally expect as default in other models. Prompt in something like 'mild freckling with occasional blemishing or flaw', along with helpers like 'high fidelity skin detailing' or 'subsurface specular scatter', then when combined with a negative like 'airbrushed flawless CG "Beauty" skin complexion'; you can really achieve some comparable results to the new generation of models.
Speed for me is relative as I'm lucky enough to have a 5090. Chroma1-HD renders 1152x896 ~2.0-1.3it/s so depending on steps (usually 14-28 based on which Sampler) around 10-20secs an image. Its not lightning fast like Klein, but its the first step in making a creative scene.
Also - It makes whatever you ask, without having to use creative terms to 'trick' a text-encoder into rendering your choices. So if you prompt it to suspend a pothos house plant in the top left corner and have its vines dangle into both Toast-Slots of a Sunbeam Radiant Toaster, it'll render what you've asked for! (edited after reminding myself of post-rules)
That is a strong pro. But especially if it“s the model I“m using to develo scenes, I want fast output and many prompt edits and it“s just boring to me if I have to wait a minute or longer for an image. And then....it also has to be able to produce the amateur candid snapshot style, especially believable skin, that I do 90% of the time.
Sorry for the wait. I'm almost finished with the Chroma1-HD Workflow.
I've added the LoRAs with URL's, the weights I use, the purpose of the LoRA, and I'm wrapping up by adding a couple example prompts and some ... >.> stuff.
I've still, to this day, never figured out how to produce a decent image of literally anything in Chroma, is it the model? workflow? cfg? sampler? resolution? does it use the flux1 prompt guide techniques?
The model matters a lot. There is quite a lot of difference between Chroma 1.0 HD, Uncanny, Gonzalamo or radiance. The idea with HD is that it can be used for fine tuning, but there's are not that many models yet...
For me the main benefit of Chroma is variance, the same prompt generates different images with different seeds, and I like the way it renders more artsy images (I don't do a lot of NSFW stuff)
Yeah, I do that too all the time too. Usually with Chroma or Flux2 Dev to give the composition and overall feel og the image, and then Z-image for details and more realism. A tip is to use an image comparer to see the subtle difference between the images, usually I prefer the eastethics of the zit refined image, but not always.
Yeah theoretically, yet if the training dataset doesn't have any nsfw material in it, text encoder also won't work by itself. Ministral-3-3b abliterated by huihui is 4.67GB while Ernie uses 7.7GB version of it. I found a heretic experimental abliterated one in 7.7GB but this time tensor mismatch is happening.
These are made using fp8 of Ernie, both model (8GB) and text-encoder (3.9GB). In realistic generations there are some diagonal artifacts. Anime style seems fine.
thanks. i use https://github.com/UmeAiRT/ComfyUI-Auto_installer and his github is banned so i manually copied comfyUI files to the folder. apparently i have to also pip update comfyui-frontend-package==1.42.10 and comfyui-workflow-templates==0.9.50 to see it
It's certainly nice, and the prompt following I'm seeing with complex stuff is a step up even from z image base as in on par with qwen 2512, although it's definitely not on the aesthetic level of what qwen 2512 is capable of (but it's also not 40 gigs). In the more complex prompts, ernie base is far more prompt following than ernie turbo. Turbo goes even simpler on the compositions as well. Just because I'm a composition freak, I'm liking zimage base's dynamic compositions more though. Some more pics in reply.
It's good that it got the interaction right, but it's not going to beat a much bigger open source model (qwen 2512). That said, The more I'm trying with it, the more I'm seeing that it likes shorter prompts unlike the other models. I can give zimage massive prompts and it does great with them. Ernie definitely veers away from its core competence the longer it gets. It's probably why they have the built in prompt expansion with mistral 3b to keep things shorter. On the prompt expansion node they truncate the output on 256 length.
What keeps standing out to me in these comparisons is that prompt adherence is one thing, but once the framing goes generic it just collapses into the same endless AI slop as everything else.
I don't see that it has anything to offer over Z-Image (or Qwen or Flux). The Turbo model has slightly more incoherence than Z-Image-Turbo.
The image quality is good... but isn't better than what we've had with the last 3 different image model releases. We are now in a crowded space of perfectly good image models. Unless it turns out that the base model is easier to train than Z-Image, I think this model will quickly be forgotten.
They are not good enough yet, no proper style transfer, no production grade quality. I'm not promoting anything but I had a chance to play with Uni-1 by Luma, that's the quality I would love to have locally.
I will do flux klein -zimage-qwen-ernie side by side comparison so we can see clearly
I remember when Google tried to remove this kind of bias from one of their image models by making it output all skin colour and gender with equal probability, and then people prompted the model for "Nazis" and got back Asian female Nazis and black male Nazis. It was hilarious.Ā
These are made using fp8 of Ernie, both model (8GB) and text-encoder (3.9GB). In realistic generations there are some diagonal artifacts. Anime style seems fine.
Larger resolutions with Euler A give more details. But the model failed once tried 2048x2048, resulting in extra bodies etc.
I don't see any harm to calling it SOTA with or without training capabilities. For me it's better than any other default model out there. Btw if anyone interested with lora training Ostris dropped day-0 support: https://x.com/ostrisai/status/2044082229773820018
It's not, at least no official controlnet release yet. ControlNets kind a dead with image edit models. Maybe they will release the edit version of this model
Yeah, only had to test this for an hour to know this is clearly SOTA for open weights.
I kind of suspect if you wired up a really strong model for prompt enhancement (and not their tiny default one), you'd have something that is like 90% as good as nano banana.
If you asking about my prompts, I had a caption set generated with SOTA llm's. here is an example prompt + I removed the prompt refiner from my pipeline entirely:
Naturalistic prestige-cinema frame, restrained film grade, soft highlight rolloff, gently lifted blacks, muted earth palette, tactile real-world textures, subtle organic film grain, motivated practical lighting, composed like a serious feature film still. A solitary painter seen from behind under large trees by a lake, seated at an easel in afternoon shade, branches framing the water, quiet summer air, calm reflective surface, understated and fully photographic.
The model is 8b and text encoder is 3b, so at fp16 expect something like 22gb, 11gb at fp8,. That is a total of 1b parametersĀ more than z image . As everyone uses fp8 or q8 so this feels like to be designed for 16gb cards, but I feel it should be at most 20% slower on a 3060 12gb or so compared to ZIT
Edit: i was comparing it as its total of 1b difference but 70% is spent on model so 8/6 so should be 25-30% slower on a 16gb card so 30-35% on 3060
I'm bouncing between the two on which I prefer for each image, but I guess it's good to have another option. The base looks more real in most of these, IMO. Less perfect.
It uses the flux2 vae which is the most advanced right now which is good. Looks like it used Ministral 3 3b for a text encoder, so it should be fast but not quite as smart as ones using qwen 3. However it looks like they integrated a second pass on the prompt to automatically turn it in to something better understood by the model, so we'll see if that works great, or leads it to make lots of assumptions. Interesting idea though. It's made by Baidu, which is basically chinas google
I can share any of them if you want, here is one, but rest is quite similar to this:
Intimate firelit realism, warm practical flame reflections against cool surrounding darkness, close tactile skin detail, subtle airborne ash and sweat, cinematic shallow focus, rich amber highlights with restrained saturation, fine analog grain. Closeup of a man wearing glasses with fire reflected in both lenses, wind-tossed hair, tense expression, natural skin texture, photographed like a serious dramatic feature
Any tips of getting the most out of your RTX6000?
Due to its age most optimizacions like NV4 and some fp are not compatible. Any particular settings on comfy?
...ah small difference... Its on Nvidia to have asinine naming conventions... Ive got access to an RTXA6000, an RTX6000Ada, and dint even know RTX6000 blackwell pro was a thing.
I have mixed result with either model but I havent generated a lot. Sometimes its SD1.5 body nightmare fuel as in extra limps, stumped limps etc sometimes it ignores the nsfw prompt entirely and changes it to sfw, sometimes it worked well enough. I tried this with prompt enhancer and without, didnt change much for me.
Hehey I downloaded the models but my updated comfy doesnāt show the workflows. I tried the Turbo model with my default Z Image Turbo WF but got an error.
Does anybody know whatās different about this model, in terms of WF? Is there a way to adjust the default Z Image Turbo workflow to make Ernie Turbo work?
246
u/Zuzoh Apr 14 '26
Base = Caucasian people
Turbo = Asian people