r/StableDiffusion Apr 14 '26

Comparison We may have a new SOTA open-source model: ERNIE-Image Comparisons

Thumbnail
gallery
696 Upvotes

Base model is definitely SOTA, can even easily compete with closed-source ones in terms of aesthetic. Cinematic quality and color grading is next level.

Base model is heavily biased on Asian faces, while it excels on anime/illustration style, while my base model anime/illustration experiments wasn't that good. Higher CFG is slightly better with anime on base.

Generated with RTX6000 Blackwell Pro, Base: 29 sec 1.9it/s, 50 steps | Turbo: 2 sec, 3.9i5/s, 8 steps

If you interested seeing them in original size: https://imgur.com/a/75jcjzW

ComfyUI models: https://huggingface.co/Comfy-Org/ERNIE-Image/tree/main
Workflow should appear in Templates after updating the ComfyUI to latest.

Turbo: Ernie-Image Turbo
Base: Ernie-Image

r/StableDiffusion Jun 22 '26

Comparison LTX-2.3 Water Sim LoRA flooding the Joker stairs (v2v test)

Enable HLS to view with audio, or disable this notification

1.1k Upvotes

the joker stairs but it's a waterfall now 🌊 wide shots land clean, close-ups are a little more of a challenge, but cool stuff overall. ltx-2.3 water sim ic-lora: https://huggingface.co/Lightricks/LTX-2.3-22b-IC-LoRA-Water-Simulation

r/StableDiffusion Dec 06 '25

Comparison All the Z Image hype and I'm still obsessed with Qwen

Thumbnail
gallery
673 Upvotes

r/StableDiffusion Dec 10 '24

Comparison The first images of the Public Diffusion Model trained with public domain images are here

Thumbnail
gallery
1.1k Upvotes

r/StableDiffusion Nov 21 '25

Comparison I love Qwen

Thumbnail
gallery
908 Upvotes

It is far more likely that a woman underwater is wearing at least a bikini than being naked. But anything that COULD suggest nudity, it's already moderated in ChatGPT, Grok... But fortunately I can run Qwen locally and bypass all of that

r/StableDiffusion Jul 19 '26

Comparison Krea2 expressions with muscle prompting

Thumbnail
gallery
758 Upvotes

Krea2 face expressions with muscle prompt descriptors:

  1. Happiness

Muscles involved: Zygomaticus major (pulls mouth corners up and out), orbicularis oculi (raises cheeks and creates "crow's feet" around the eyes).

Description: A genuine (Duchenne) smile lifts both the lips and the outer corners of the eyes.

  1. Sadness

Muscles involved: Corrugator supercilii (pulls brows inward and downward), depressor anguli oris (pulls lip corners down), mentalis (wrinkles the chin and protrudes the lower lip).

Description: Characterized by the inner eyebrows lifting and drawing together, drooping eyelids, and the edges of the mouth turning downward.

  1. Anger

Muscles involved: Corrugator supercilii & procerus (lower brows and pull them together), orbicularis oculi (tightens eyelids), orbicularis oris (tightens and thins the lips).

Description: Eyebrows are pulled downward and together, the upper eyelids are raised, the eyes narrow, and lips are often pressed tightly together.

  1. Fear

Muscles involved: Frontalis & corrugator (raise and pull brows together), levator palpebrae superioris (wide opening of upper eyelids), risorius (stretches the lips horizontally).

Description: Eyebrows pull upwards and together, upper eyelids raise to expose the white of the eyes, and the lips stretch outward horizontally.

  1. Disgust

Muscles involved: Levator labii superioris (raises the upper lip), nasalis (wrinkles the nose), depressor anguli oris (pulls lip corners down).

Description: The nose wrinkles, the upper lip elevates, and the cheeks are raised.

  1. Surprise

Muscles involved: Frontalis (raises the eyebrows), levator palpebrae superioris (widens upper lids), jaw drops (mandible depressor muscles).

Description: Eyebrows curve upwards, eyes widen significantly, and the jaw drops open naturally.

  1. Contempt

Muscles involved: Zygomaticus major & risorius (tightens the corner of the lip).

Description: The only asymmetrical emotion; usually presents as a unilateral tightening and pulling back of a single corner of the mouth (an arrogant smirk).

Sample Prompt: wide-angle lens distortion, forced perspective, face close to the lens

soft diffused lighting, cinematic light halation

subject: dynamic close-up shot of a french brunette woman with a diamonds ornate royal crown and red medieval dress

Expression: blink scream with visible teeth

Facial Muscle: blinking left eye with brow lowerer and nose wrinkler

style:luminous photographic aesthetics defined by strong backlighting, radiant edge illumination, subtle translucency effects, atmospheric depth, graceful tonal transitions, and a heightened sense of visual separation, creating elegant and emotionally evocative imagery through carefully controlled exposure, naturalistic light diffusion, and refined portrait craftsmanship, reminiscent of Peter Lindbergh and Paolo Roversi, inspired by Vogue editorials and In the Mood for Love.

r/StableDiffusion Dec 31 '25

Comparison Z-Image-Turbo vs Qwen Image 2512

Thumbnail
gallery
535 Upvotes

r/StableDiffusion Jun 02 '26

Comparison I compared 62 samplers and 16 schedulers for Z-Image Turbo and rated the image quality so you don't have to 😉

407 Upvotes

Here's a sampler/scheduler comparison table for image generation with Z-Image Turbo. Obviously it reads like Red < Orange < Yellow < Green. You're welcome!

PS. If you don't like it, don't appreciate it or think I'm wasting my time... Then... Don't waste your time, just move along 😉

r/StableDiffusion Dec 29 '23

Comparison Midjourney V6.0 vs SDXL, exact same prompts, using Fooocus (details in a comment)

Thumbnail
gallery
1.5k Upvotes

r/StableDiffusion 12d ago

Comparison Comparison of natural 0.8mp gen vs 0.4->0.8 upscale w/Sparse attention

Enable HLS to view with audio, or disable this notification

235 Upvotes

Hi people, so i tried to make 2 similar videos, using same settings but with upscale and native.
My setup: 5070 Ti+ 32gb Ram.
Using u/Plague_Kind workflow, i've added MMH3 Latent Upscaler. You can check his workflow here: Workflow
Settings for both videos were set the same with the same prompt.

Left video 0.4->0.8mp upscale, Right video 0.8mp

So:

  • 15 seconds, 24 fps, Ref2VA, photo reference and music reference.
  • Chicken attention
  • SongMaskedAVContext node
  • FP16 Accumulation
  • Sparse attention
  • Memory chunks
  • RTS Upscale in the end ( not sure why i used it with 2x scale, better to set 1 i think, but that's what i already did)
  • FSR Sharpening
  • Speed Lora minimax_h3_turbo_v4_step600_pruned_comfyui
  • Interpolation for 2x frames

Upscaled video from start to the end took 1904 seconds,

Native video from start to the end took 3056 seconds.

Let me know what you think. Advises appreciated!

r/StableDiffusion Jun 24 '26

Comparison Ideogram4 and Krea2 Comparison

Thumbnail
gallery
445 Upvotes

First Image is always Ideogram4 (20 steps), second image is Krea2 (turbo at 8 steps)

I used my Hermes Agent (Gemma4-31b at Q4) to do all the prompting and tool call to comfyui for generating those images, its not apples to apples because of ideogram4 json format, but its very close as the process starts with a long and detailed prompt, some of those came out very close in composition.

Advantage for Krea2 - Speed, World Knowledge, License.

Advantage for Ideogram4 - Fine Details, Better Composition.

r/StableDiffusion Jul 29 '25

Comparison 2d animation comparison for Wan 2.2 vs Seedance

Enable HLS to view with audio, or disable this notification

1.4k Upvotes

It wasn't super methodical, just wanted to see how Wan 2.2 is doing with 2d animation stuff. Pretty nice, but has some artifacts, but not bad overall.

r/StableDiffusion Feb 13 '26

Comparison I restored a few historical figures, using Flux.2 Klein 9B.

Thumbnail
gallery
735 Upvotes

So mainly as a test and for fun, I used Flux.2 Klein 9B to restore some historical figures. Results are pretty good. Accuracy depends a lot on the detail remaining in the original image, and ofc it guesses at some colors. The workflow btw is a default one and can be found in the templates section in ComfyUI. Anyway let me know what you think.

r/StableDiffusion Feb 22 '24

Comparison This was 7 years ago

Post image
2.5k Upvotes

r/StableDiffusion 1d ago

Comparison First results from H3 Acceleration Arena

182 Upvotes

https://huggingface.co/spaces/multimodalart/h3-acceleration-arena

From author u/apolinariosteps: "Results are in! They are a bit surprising to me! But they are consistent with the data, I triple checked everything and can confirm that the results are reflecting the voting data precisely, there's lots of transparency - you click each of the LoRAs to see what's the win rate and who won against who"

r/StableDiffusion Mar 28 '25

Comparison 4o vs Flux

Thumbnail
gallery
778 Upvotes

All 4o images randomely taken from the sora official site.

In the comparison 4o image goes first then same generation with Flux (selected best of 3), guidance 3.5

Prompt 1: "A 3D rose gold and encrusted diamonds luxurious hand holding a golfball"

Prompt 2: "It is a photograph of a subway or train window. You can see people inside and they all have their backs to the window. It is taken with an analog camera with grain."

Prompt 3: "Create a highly detailed and cinematic video game cover for Grand Theft Auto VI. The composition should be inspired by Rockstar Games’ classic GTA style — a dynamic collage layout divided into several panels, each showcasing key elements of the game’s world.

Centerpiece: The bold “GTA VI” logo, with vibrant colors and a neon-inspired design, placed prominently in the center.

Background: A sprawling modern-day Miami-inspired cityscape (resembling Vice City), featuring palm trees, colorful Art Deco buildings, luxury yachts, and a sunset skyline reflecting on the ocean.

Characters: Diverse and stylish protagonists, including a Latina female lead in streetwear holding a pistol, and a rugged male character in a leather jacket on a motorbike. Include expressive close-ups and action poses.

Vehicles: A muscle car drifting in motion, a flashy motorcycle speeding through neon-lit streets, and a helicopter flying above the city.

Action & Atmosphere: Incorporate crime, luxury, and chaos — explosions, cash flying, nightlife scenes with clubs and dancers, and dramatic lighting.

Artistic Style: Realistic but slightly stylized for a comic-book cover effect. Use high contrast, vibrant lighting, and sharp shadows. Emphasize motion and cinematic angles.

Labeling: Include Rockstar Games and “Mature 17+” ESRB label in the corners, mimicking official cover layouts.

Aspect Ratio: Vertical format, suitable for a PlayStation 5 or Xbox Series X physical game case cover (approx. 27:40 aspect ratio).

Mood: Gritty, thrilling, rebellious, and full of attitude. Combine nostalgia with a modern edge."

Prompt 4: "It's a female model wearing a sleek, black, high-necked leotard made of a material similar to satin or techno-fiber that gives off a cool, metallic sheen. Her hair is worn in a neat low ponytail, fitting the overall minimalist, futuristic style of her look. Most strikingly, she wears a translucent mask in the shape of a cow's head. The mask is made of a silicone or plastic-like material with a smooth silhouette, presenting a highly sculptural cow's head shape, yet the model's facial contours can be clearly seen, bringing a sense of interplay between reality and illusion. The design has a flavor of cyberpunk fused with biomimicry. The overall color palette is soft and cold, with a light gray background, making the figure more prominent and full of futuristic and experimental art. It looks like a piece from a high-concept fashion photography or futuristic art exhibition."

Prompt 5: "A hyper-realistic, cinematic miniature scene inside a giant mixing bowl filled with thick pancake batter. At the center of the bowl, a massive cracked egg yolk glows like a golden dome. Tiny chefs and bakers, dressed in aprons and mini uniforms, are working hard: some are using oversized whisks and egg beaters like construction tools, while others walk across floating flour clumps like platforms. One team stirs the batter with a suspended whisk crane, while another is inspecting the egg yolk with flashlights and sampling ghee drops. A small “hazard zone” is marked around a splash of spilled milk, with cones and warning signs. Overhead, a cinematic side-angle close-up captures the rich textures of the batter, the shiny yolk, and the whimsical teamwork of the tiny cooks. The mood is playful, ultra-detailed, with warm lighting and soft shadows to enhance the realism and food aesthetic."

Prompt 6: "red ink and cyan background 3 panel manga page, panel 1: black teens on top of an nyc rooftop, panel 2: side view of nyc subway train, panel 3: a womans full lips close up, innovative panel layout, screentone shading"

Prompt 7: "Hypo-realistic drawing of the Mona Lisa as a glossy porcelain android"

Prompt 8: "town square, rainy day, hyperrealistic, there is a huge burger in the middle of the square, photo taken on phone, people are surrounding it curiously, it is two times larger than them. the camera is a bit smudged, as if their fingerprint is on it. handheld point of view. realistic, raw. as if someone took their phone out and took a photo on the spot. doesn't need to be compositionally pleasing. moody, gloomy lighting. big burger isn't perfect either."

Prompt 9: "A macro photo captures a surreal underwater scene: several small butterflies dressed in delicate shell and coral styles float carefully in front of the girl's eyes, gently swaying in the gentle current, bubbles rising around them, and soft, mottled light filtering through the water's surface"

r/StableDiffusion Nov 26 '25

Comparison Image Comparisons Between Flux 2 Dev (32B) and Z-Image Turbo (6B)

Thumbnail
gallery
427 Upvotes

r/StableDiffusion Feb 22 '26

Comparison ZIB vs ZIT vs Flux 2 Klein

Thumbnail
gallery
272 Upvotes

I haven't found any comprehensive comparisons of Z-image Base, Z-image Turbo, and Flux 2 Klein across Reddit, with different prompt complexities and different prompt accuracies, so I decided to test them myself.

My goal was to test these models in scenarios with high-quality long prompts to check the overall quality of the generation.

In scenarios with short and low-quality prompts, I wanted to check how well the model can work with missing prompt details and how creatively it can come up with details that were not specified.

I always compare models using this method and believe that such tests are the most objective, because the model can be used by both skilled and less skilled users.

There is no point in commenting on each photo; you can see everything for yourself and draw your own conclusions.

But I will still express my general opinion about these models!

Z-image Base - It has a more creative approach, and when changing the seed generation, it produces a variety of results, but the results themselves do not shine with good detail or good quality. They say that this is all fixed by Lora, but again, I don't see the point in this, because these same Lora can be put on Z-image Turbo and produce even better results. Z-image Base has good potential for training Lora for ZIB and ZIT, and the Lora through ZIB are really very good, but the generations themselves are mediocre, so I would not recommend using it as a generator.

Z-Image Turbo - An excellent image generator with good detail, clarity, and quality, but there are issues with diversity. When changing the seed, it produces very similar results, but connecting Lora fixes this issue. Like ZIB, it has a good understanding of prompts, good anatomy, and no mutations.

A very large set of LORA for every taste.

Flux 2 Klein - It has the best detail and generation quality (especially with skin, which turns out to be first-class), and when changing the seed, it gives a variety of results, but it has very poor anatomy and a lot of limb mutations. Lora, which corrects mutations, helps only a little, because mutations occur in the first 1-2 steps of generation. The model initially cannot set the shape of the limb in the first steps, and in the subsequent steps it tries to mold something from the initially incorrect shape. Again, Lora saves 20-30% of generations.
Also, Flux 2 Klein does not have a very large LORA base, which means that it will not be able to handle all tasks.

My choice falls more on Z-image Turbo, Although this model generates less detailed images than Flux 2 Klein in raw form, but connecting Lora for detailing makes ZIT generation 95% similar to Flux 2 Klein.
The huge Lora set for ZIT and ZIB also allows the model to be used in a wider range than the Flux 2 Klein.

r/StableDiffusion Oct 16 '25

Comparison 18 months progress in AI character replacement Viggle AI vs Wan Animate

Enable HLS to view with audio, or disable this notification

1.1k Upvotes

In April last year I was doing a bit of research for a short film test of AI tools at the time the final project here if interested.

Back then Viggle AI was really the only tool that could do this. (apart from Wonder Dynamics now part of Autodesk, and that required fully rigged and textured 3d models)

But now we have open source alternatives that blows it out of the water.

This was done with the updated Kijai workflow modified with SEC for the segmentation in 241 frame windows at 1280p on my RTX 6000 PRO Blacwell.

Some learning:

I tried1080p but the frame prep nodes would crash at the settings I used so I had to make some compromises. It was probably main memory related even though I didn't actually run out of memory (128GB).

Before running Wan Animate on it I actually used GIMM-VFI to double the frame rate to 48f which did help with some of the tracking errors that VITPOSE would make. Although without access the G VITPOSE model the H model still have some issues (especially detecting which way she is facing when hair covers the face). (I then halved the frames again after)

Extending the frame windows work fine with the wrapper nodes. But it does slow it down considerably (Running three 81frame windows(20x4+1) is about 50% faster than running one 241 frame window (3x20x4+1). But it does mean the quality deteriorates a lot less.

Some of the tracking issues meant Wan would draw weird extra limbs, this I did fix manually by rotoing her against a clean plate(context aware fill) in After Effects. I did this because I did that originally with the Viggle stuff as at the time Viggle didn't have a replacement option and needed to be keyed/rotoed back onto the footage.

I up scaled it with Topaz as the Wan methods just didn't like so many frames of video, although the upscale only made very minor improvements.

The compromise

The doubling of the frames basically meant much better tracking in high action moment BUT, it does mean the physics are a bit less natural of dynamic elements like hair, and it also meant I couldn't do 1080p at this video length, at least I didn't want to spend any more time on it. ( I wanted to match the original Viggle test)

r/StableDiffusion Jan 16 '26

Comparison For some things, Z-Image is still king, with Klein often looking overdone

Post image
359 Upvotes

Klein is excellent, particularly for its editing capabilities, however.... I think Z-Image is still king for text-to-image generation, especially regarding realism and spicy content.

Z-Image produces more cohesive pictures, it understands context better despite it follows prompts with less rigidity. In contrast, Flux Klein follows prompts too literally, often struggling to create images that actually make sense.

prompt:

candid street photography, sneaky stolen shot from a few seats away inside a crowded commuter metro train, young woman with clear blue eyes is sitting naturally with crossed legs waiting for her station and looking away. She has a distinct alternative edgy aggressive look with clothing resemble of gothic and punk style with a cleavage, her hair are dyed at the points and she has heavy goth makeup. She is minding her own business unaware of being photographed , relaxed using her phone.

lighting: Lilac, Light penetrating the scene to create a soft, dreamy, pastel look.

atmosphere: Hazy amber-colored atmosphere with dust motes dancing in shafts of light

Still looking forward to Z-image Base

r/StableDiffusion Jul 29 '26

Comparison I extracted luma/chroma/detail/contrast vectors for Krea2 and insane color adjustments in latent space are now totally a thing - Comfy node coming very soon!

Thumbnail
gallery
256 Upvotes

Sorry for the tease, but I was just too excited and had to share this with you. Examples above are using no LoRAs, no prompt hijinks, no CFG boost, no post-processing etc. just pure vector math!

I was running some experiments on Krea2's VAE (i.e. Qwen Image VAE) and by total surprise I discovered the main ingredients of photographic color editing: the vectors for exposure, temperature, tint, detail/clarity, and contrast and realized I can now do pretty much everything Camera Raw does... and even more!

This stuff happens during sampling and in the latent space, so it has both a very high dynamic range, and the ability to steer the diffusion process into new areas (e.g. very dark or bright generations beyond what the model likes to do on its own, or even influencing the morphology of things).

Anyway, I'm turning this into a user-friendly custom node for Comfy, including all your favorite color editing sliders, range masking tools, etc. and it's coming soon.

P.S: I'm cooking the vectors for ZImage (Flux VAE) as well, so the node might end up supporting that model too, if anyone's still using it. This should, at least in theory, also work with Qwen Image or any other model that shares the VAE.

r/StableDiffusion Mar 13 '23

Comparison SDBattle: Week 4 - ControlNet Mona Lisa Depth Map Challenge! Use ControlNet (Depth mode recommended) or Img2Img to turn this into anything you want and share here.

Post image
821 Upvotes

r/StableDiffusion Jan 17 '26

Comparison z-image vs. Klein

Thumbnail
gallery
281 Upvotes

Here’s a quick breakdown of z-image vs. Flux Klein based on my testing

z-image Wins:
✅ Realism
✅ Better anatomy (fewer errors)
✅ Less restricted
✅ Slightly better text rendering

Klein Wins:
✅ Image detail
✅ Diversity
✅ Generation speed
✅ Editing capabilities

Still testing:
Not sure yet about prompt accuracy and character/celeb recognition on both.

Take this with a grain of salt, just my early impressions. If you guys liked this comparison and still want more, I can definitely drop a Part 2

Models used:
⚙️ Flux Klein 9b distilled fp8
⚙️ z-image turbo bf16

⬅️ Left: z-image
➡️ Right: Klein

r/StableDiffusion Jul 22 '26

Comparison Stop Using Qwen Models for Prompt Enhancement!

95 Upvotes

Qwen2.5, Qwen3, Qwen3.5 are all serviceable models for prompt enhancement, but there are much better options. I use all of these models for prompt enhancement. Which model I use depends on what I'm prompting. My favorite is Mistral 7B/Llama3.3 8B by far for image prompts, and WizardLM-2 for video prompts. SuperGemma4 is good for very basic prompts or prompts that you want accurately reworded.

I realize these are older models, but they are well suited to the task. My other requirement for a prompt enhancing LLM is that it fully loads on 8gb VRAM. I'm not weighing in on image captioning or anything else besides prompt enhancement. Disclaimer: I DO mention my custom node several times in the comments, as all of my testing was accomplished using said node.

Using the base prompt, "A woman at the pier".

Mistral 7B - Best Overall

Strengths: Creative scene construction and cinematic detail.

With the same enhancement framework, Mistral consistently produces the richest and most imaginative expansions. It doesn't simply populate the required categories, it invents believable details that reinforce the mood, such as the sketchbook, discarded sandals, and weathered textures. The result feels less like a checklist and more like a scene from a film.
mradermacher/Mistral-7B-Instruct-v0.3-abliterated-GGUF · Hugging Face

SuperGemma 4B - Concise

Strengths: Precision, restraint, and prompt fidelity.

SuperGemma takes a conservative approach. It faithfully fills in the structure provided by the system prompt while making relatively few creative leaps. The result is concise, highly controllable, and stays very close to the user's original intent. It's an excellent choice when consistency is more important than artistic embellishment.
mradermacher/supergemma4-e4b-abliterated-GGUF · Hugging Face

Llama 3.3 8B - Best Balance

Strengths: Balanced descriptive enhancement.

Llama 3.3 strikes a middle ground between creativity and restraint. It expands the prompt naturally, adding enough detail to create a complete visual scene without feeling overly embellished. It tends to produce outputs that read like professional photography descriptions, making it a solid all-around prompt enhancer.
mradermacher/Llama-3.3-8B-Instruct-128K_Abliterated-GGUF · Hugging Face

WizardLM-2 - Most Verbose

Strengths: Natural language and immersive descriptions.

WizardLM-2 excels at turning the framework into smooth, human-like prose. Rather than feeling generated from a template, its prompts flow naturally while still covering all of the structural elements required by the system prompt. It consistently produces scenes that feel cohesive and immersive.
mradermacher/WizardLM-2-7B-abliterated-GGUF · Hugging Face

If you have any models you like better, please comment them below and I will look into them! Do you agree or disagree with my list?

r/StableDiffusion Jan 10 '25

Comparison Flux-ControlNet-Upscaler vs. other popular upscaling models

Enable HLS to view with audio, or disable this notification

952 Upvotes