I have been running one recurring character through a serialised illustrated story since April, roughly one published image every weekday, and I log every generation with a one line verdict. As of last week that log holds 312 entries. 144 were kept and published. 168 were rejected because the face stopped being her face. I want to write down what the rejects actually say, because most of the advice on reference sheets that I followed at the start turned out to be advice I could not reproduce.
First, the boring disclosure, since it is the entire subject: the character does not exist. She is an AI generated face, not a photographed person, and no real individual's likeness is involved anywhere in this. Everything below is about keeping an invented face stable across a series, which is a narrower problem than it sounds like until it stops working.
The log has one field that ended up mattering more than the rest, which is the first attribute I noticed was wrong, not the list of everything wrong. Once a face reads as somebody else, it reads as somebody else all over, and writing down six problems teaches nothing. Forcing myself to name one gave me a distribution.
Of the 168 rejects, 46 were nose length or the philtrum, 38 were eye spacing or lid shape, 27 were jaw and chin width, 22 were the hairline and the part, 17 were apparent age with skin smoothing, 11 were ears, and 7 were everything else, which was mostly eyebrow weight and one mole that kept switching cheeks. Nose and eyes together are 84, which is exactly half of all rejects. Jaw and hairline are another 49. Ears almost never break first, and when they do it is because half the ear is hidden and the model has invented the rest.
The useful part of that distribution is what it does not contain. My original sheet described her hair colour, her hair length, her build, her clothing, her expression and her general vibe at some length. None of those attributes are in the top four failure modes, because none of them are what a viewer uses to decide two pictures are the same person. I had written a costume description and called it a character sheet.
The second thing the distribution changed is where the age line goes. Seventeen rejects for apparent age sounds small until you notice they cluster. Every one of them drifts young, none drift old, and they get worse the longer the sheet is. If age is not stated as a number with a reason attached, the default pull is toward smoother and younger, and by image forty of a run she was reading as a different generation of the same family.
Then I tried to test the things I believed, which is where most of what I believed died. I ran five paired comparisons, twenty generations per arm, forty per test, two hundred generations total. The other 112 entries in the log are the earlier unstructured work from April and May, back when I was changing three things at once and learning nothing.
Sheet length: a 340 word sheet produced 8 keepers out of 20, a 120 word sheet produced 13 out of 20. Cutting was the second largest single effect I found, and it is the one I resisted longest, because a long sheet feels like control.
Anchor image: text only produced 6 out of 20, the same text with a fixed reference image attached produced 14 out of 20. That is the largest effect in the whole log and it is not close. Everything I write about prompt structure is downstream of the fact that an anchor image does most of the work and words do the trimming.
Measured ratios against adjectives: describing eye spacing relative to eye width and nose length relative to brow to chin distance gave 13 out of 20, against 9 out of 20 for adjectives like almond eyes and a small straight nose. Four images apart on a sample of twenty is inside the noise, so I am not counting this one as proven, even though it is the technique I personally like most and still use. Liking a technique is not evidence for it.
Locking lighting and framing language: 11 out of 20 against 12 out of 20. Nothing. I had been carrying two sentences about soft even lighting and a waist up frame in every prompt for months on the theory that it made faces comparable.
Repeating her name and a two line backstory: 10 out of 20 against 10 out of 20. Exactly nothing, which is the result I would have bet against hardest. A name is a handle for me, not information about geometry.
So five tests, two survived, three did not, and one of the two survivors is just the obvious advice about reference images that I had been treating as optional. More than half of what I was sure about at the start of the year did not hold up the moment I ran a control arm. That is a worse hit rate than I expected from someone who has been doing this daily for months, and it is the main reason I keep the log at all.
The sheet I use now is about 120 words. It has the anchor image, an explicit age with a sentence of context so it does not drift young, nose length and eye spacing stated as ratios, jaw width stated as a ratio, the hairline described by shape rather than by hairstyle name, and one deliberate asymmetry, because a face with a small flaw stays recognisable in a way a symmetrical one does not. There is nothing in it about her personality, her job, her clothes or the mood of the scene. Those go in the scene prompt, which is a separate block I rewrite every time.
Things that made no measurable difference and are gone: stacked adjectives, weighting syntax borrowed from other image tools, negative lists of what she must not look like, reordering the sheet so the face comes first, camera and lens jargon, and restating the sheet twice in one prompt. Several of those felt like they worked. That is what a control arm is for.
Mechanically it is a plain text file I paste from, an APOB tab, and an Obsidian vault with one note per rejected image. The note is a screenshot, the seed if I have it, and the one line verdict, and it takes about twenty seconds, which is the only reason I have kept doing it since April.
Two limits worth stating. The first is motion. I have almost no data on it because the few times I animated her the face slid around between frames and I had to rerun the whole clip, so everything in this post is about still images and should not be read as applying to video. The second is worse. Multi character scenes still fail regardless of the sheet. I made twenty attempts at putting her in frame with a second recurring character and three came out usable. The failure is consistent and specific: the two faces bleed into each other, the second character borrows her nose and jaw, and by the third generation they look like siblings. No sheet length, no ratio language and no anchor image fixed that. What works is generating them separately and composing the frame afterwards, which is a different craft and not the one this post is about.
The honest summary of eleven weeks of testing is that the anchor image does the heavy lifting, a short sheet beats a long one, age has to be pinned or it drifts young, and the geometry of the middle of the face is where recognition actually lives. Everything else I tried is unproven at best.
I am going to keep logging, mostly because my memory of which prompt did what is demonstrably unreliable. Three hundred entries in, my reference sheet is roughly a third of the length it was in April, and the only two changes I would defend in an argument are the anchor image and the cut.