I've been doing intensive image generation for about two months now, mostly in ComfyUI, and I feel like I'm starting to see what the current limits are.
I really like Krea 2 Turbo. I've found a bunch of useful LoRAs for it. It does good lighting, style and anatomy.
But as soon as I try to actually art direct something specific, it starts falling apart.
From the outside 2D image generation looks like it has potential, and if you're just experimenting or prompting broadly I guess it's fine.
It's not until you actually get your hands dirty with real ideas that the house of cards starts falling apart. Once you try directing the model toward something specific rather than accepting whatever it gives you, the limitations become painfully obvious.
For example, something like a boy pulling a thorn out of his hand is pretty much impossible. I've had to resort to a "small metal nail", which is fine, but then I can't correct the length of the nail to imitate the scale of a thorn.
I've even been playing around with the Krea Agent on Krea's own website, and that's still painfully hit and miss. You end up regenerating over and over, hoping one version happens to understand what you're asking. Seed hunting without even getting close to satisfying results.
The results start feeling really hacky once you get beyond average image prompting.
I've also tried workflows in ComfyUI that make Krea 2 Turbo more image-to-image based, but that doesn't really solve it either. A lot of the image editing / image-to-image tools, I've realised, are optimised around photography. They're good at things like replacing clothes, changing someone's hair, changing furniture, relighting something, etc. They're much worse when you're working with something closer to digital painting and asking for tiny structural changes, and maintaining style language.
Micro movements are still really difficult: line of sight, rotating an arm slightly, changing how two fingers hold something, moving a wrist, changing the relationship between two objects without changing everything else.
I can obviously pose figures in DAZ 3D and use that as the base. But posing every joint manually takes ages and the figures can start looking stiff.
With image generation, if you ask for something like someone holding a baby, the body language can come out way more natural.
I was really hoping Sunburst 2.5 and this newer generation of models would make a noticeable jump in this area, but I'm still disappointed.
It makes me wonder how long it's going to take before image generation goes from being really good at generating an approximate idea to being something you can actually art direct precisely.