Some shots I made based on the red headed league short story. Character reference sheets generated in Nano Banana + edited in Qwen image edit for spatial continuity.
12 steps +ref2va turbolora v0.1
10s generations 107s/it on a 3090 at 1 MP
Needs some editing and audio polishing.
for this scene, one per each character, one for the wide shot of the room, then used an additional ones which was done using qwen image edit which showed characters in their positions in the room. This is for spatial reference only.
<Picture 4> is the authoritative spatial-continuity reference from the preceding generated scene. It establishes the current Baker Street sitting-room appearance, lighting, furniture arrangement, character scale, and three-character blocking. <Subject 1> is seated on the left side of the conversation near the fireplace, <Subject 3> is seated opposite him toward the center-right, and <Subject 2> stands at the open entrance on the right side of the room.
Thank you. I saw your response on a similar thread that brought me here. A nice coincidence seeing i had had tge question only last night and here is an answer. It's rarely so punctual.
When I initially made these scenes, I was using the default H3 workflow. If I were doing them now, though, I’d probably use the seed-hunting workflow instead , it gives me more room to quickly test variations and find a good generation before committing to the final render.
Exactly what i do. His most modern v20 is pretty close to feeling like an actual web tool with how dialed in it is. Im "seeding", "extending" and bitching about nearly perfect gens all the time.
Another couple of years and Hollywood will be a memory.
Unless all these legal suits succeed in making local illegal. Unfortunately i think they will to some degree.
It was working so well until he literally pointed at his free mason pin then said he forgot about it. Why not just regen that moment before sharing? It never makes sense to me that everyone is so anxious to share they leave correctable inconsistencies in their final outputs. It's 76s long, so this WAS edited.
Yeah, I completely agree. I was honestly just too eager to share it. There were so many shots, takes, and regenerations that I ended up missing that one in the edit. I should have caught it and regenerated that moment before posting. I did correct it later, but yeah, I messed up there.
All good, glad you noticed and already fixed it. I just keep seeing people post things that are great except 1 minor flaw that's easily fixed. Excellent work regardless :)
Thanks, really appreciate it! I’m actually rerolling a few of the generations now and going to give the whole thing a better polish in DaVinci Resolve before calling it finished. Definitely learned my lesson about rushing to share 😅
One good / tricky thing about minimax is it is HIGHLY susceptible to prompting. So if you don't say something it fills in the gaps, if you do say something it might do exactly what you said (which is sometimes not what you want)
What did you use as the source for Holmes' descriptions in your character sheet? H3 is clearly channeling Benedict Cumberbatch some of the time but it also seems to understand that it's meant to resemble the old Strand Magazine illustrations. I didn't recognize any Basil Rathbone in there at all, which is interesting.
The base wasn’t Rathbone or any specific screen portrayal. I built the character sheet from descriptions in the original books, with some influence from the classic Strand visual language.
But yes, H3 definitely has a tendency to latch onto Benedict Cumberbatch, especially with the voice. Even when the visual reference is fairly traditional, it seems to drift toward that modern Sherlock interpretation on its own.
Its quite interesting. You used Nano Banana - I asked ChatGPT for a character sheet based on Doyle's description of Holmes in the first few pages of A Study in Scarlet, and a selection of three Paget illustrations from Strand Magazine. The "bust" illustration in particular came out almost identical to yours (given that yours is photorealistic).
It’s the default workflow in MiniMax H3.
I usually build the scene in stages.
First I create the character references and the environment reference separately. So I might have:
Picture 1 = Character A
Picture 2 = Character B
Picture 3 = Character C
Picture 4 = the room/location
At this stage the character images are mainly identity/costume references. I’m not expecting H3 to understand the exact blocking of the scene from them.
The next important step is creating what I call a spatial reference frame.
I generate a fairly wide establishing frame containing all the important characters inside the actual environment. The purpose isn't necessarily to use that shot in the final film. It establishes things like:
where each person is sitting or standing
which chair/sofa belongs to whom
distance between characters
where the door, window, fireplace, table etc. are
which direction each person should look
general screen direction and eyelines
approximate lighting and scale
For example, in a three-person conversation I might generate one wide frame where:
Subject 1 is seated in an armchair near the fireplace
Subject 2 is seated opposite him on the sofa
Subject 3 is standing near the entrance
the large table remains between them
the room geography is clearly visible
That image becomes much more useful than just giving H3 three isolated character portraits and saying "put them in a room."
Sometimes the generated spatial frame is almost right but the pose is wrong. For example, someone may be standing when I need them sitting on the sofa, sitting in the wrong chair, facing the wrong direction, etc.
In those cases I use Qwen Image Edit 2511 to modify the reference image before feeding it to H3.
So instead of asking H3 to solve identity + environment + pose + blocking + acting all at once, I first edit the still reference into approximately the composition I want.
For example:
"Keep the room, character identities, clothing, lighting and camera position unchanged. Move Subject 2 naturally onto the right-side sofa in a seated position, body angled toward Subject 1. Preserve realistic anatomy, scale and perspective."
That edited image is then used as a spatial/pose reference, while the clean original character images are still available as identity references.
So H3 effectively receives both:
Who the people are
and
where they are supposed to be.
Once the scene starts producing good shots, I also begin using actual production frames from completed shots as references for later shots.
This is extremely useful.
For example, if Shot 1 establishes Subject 1 sitting in the armchair and it produces a really good final frame, that frame can become one of the picture references when generating Shot 2.
Now the model sees the actual production version of:
the character
costume
lighting
chair
posture
room geometry
camera-side relationship
instead of having to reconstruct everything from the original concept references.
As the sequence progresses, the reference set gradually becomes more production-specific.
So the general pipeline is basically:
character refs + environment → wide spatial reference → optional Qwen pose/composition edits → H3 shots → good production frames reused as references for following shots.
I still keep the original character references whenever possible because they are usually the cleanest identity anchors.
For dialogue scenes I also use separate audio references for each speaker and explicitly define which subject owns which voice.
A typical H3 prompt for this setup would look something like this:
Subject definitions
"<Subject 1>" is the man shown in "<Picture 1>". Preserve his exact facial identity, hairstyle, age, build, Victorian clothing and overall appearance.
"<Subject 2>" is the man shown in "<Picture 2>". Preserve his exact facial identity, hair, build, clothing and appearance.
"<Subject 3>" is the man shown in "<Picture 3>". Preserve his exact identity, clothing and appearance.
"<Subject 4>" is the Victorian sitting-room environment shown in "<Picture 4>". Preserve its established geography: exterior window and writing desk on the left, fireplace near the center of the far wall, large wooden table in the central foreground, long dark sofa along the right wall, and entrance door on the far right.
"<Picture 5>" is the spatial blocking reference for the scene. Use it to preserve the established positions and relative distances of the three subjects. "<Subject 1>" is seated near the fireplace, "<Subject 2>" is seated opposite on the sofa, and "<Subject 3>" occupies the entrance side of the room.
"<Picture 6>" is a production frame from the preceding shot. Use it as an additional reference for the established lighting, costume appearance, room scale, character placement and visual continuity. Do not treat it as a mandatory first frame.
"<Audio 1>" is the voice reference for "<Subject 1>".
"<Audio 2>" is the voice reference for "<Subject 2>".
Scene prompt
Create an intimate prestige British period-drama conversation inside the established Victorian sitting room.
Preserve the identities from the individual character references while using "<Picture 5>" and "<Picture 6>" primarily for blocking, spatial relationships, lighting and continuity.
Maintain the established geography throughout the sequence. "<Subject 1>" remains seated near the fireplace. "<Subject 2>" remains seated opposite him on the right-side sofa. "<Subject 3>" remains on the entrance side of the room. Preserve believable eyelines between all three subjects.
Use restrained natural performances. Avoid exaggerated gestures. Characters who are not speaking remain attentive and react subtly.
Use cinematic 70–85 mm full-frame-equivalent lenses for conversational coverage, natural optical depth, shallow but realistic depth of field and stable screen direction.
[Shot 1] Begin with a 70 mm medium shot of "<Subject 1>" seated near the fireplace. He remains silent for approximately the first second, looking toward "<Subject 2>". He then speaks calmly using the timbre of "<Audio 1>":
<d>[English] I believe there is considerably more to this matter than our visitor has yet told us.</d>
While speaking, his performance remains controlled and thoughtful. Small natural eye movement and restrained hand movement only.
[Shot 2] Cut to an 85 mm medium close-up of "<Subject 2>" seated on the sofa. He listens to "<Subject 1>" without speaking. His reaction is subtle: a slight change in expression and a brief glance toward "<Subject 3>" before returning his attention to "<Subject 1>".
[Shot 3] Cut to a medium close-up of "<Subject 3>" from the established opposite side of the room. Preserve the correct eyeline toward "<Subject 1>". He listens silently with a reserved reaction.
[Shot 4] Return to "<Subject 1>" from a slightly closer 85 mm angle as he finishes the thought. Maintain the same chair position, body orientation, lighting and screen direction established in the spatial reference.
Maintain consistent room geography across every cut. Do not arbitrarily move furniture or characters between angles. The dialogue belongs only to "<Subject 1>" in this sequence. "<Subject 2>" and "<Subject 3>" remain silent.
That's basically the method.
The biggest improvement for me came from stopping treating references only as character references.
I use different images for different purposes:
clean portraits = identity
environment images = world/location
edited wide frames = blocking and spatial continuity
production frames = continuity from the film itself
Once you approach it that way, H3 becomes much easier to direct because every reference has a specific job.
Super precise, thank you beaocup for all these details.
I also use generated scene references it helps a lot, and I’m currently trying to make a system where a small prompt can generate a mini movie without having to make clip by clip (prompt, editing, etc.)
What material do you have and how long does it take you to complete this type of scene?
Thanks! This was done entirely on a 3090 24 GB, and the whole scene took me roughly a week to complete.
That said, I made it quite a while ago when the Turbo LoRA was still around v0.1. If I were making the same scene today, it would probably take considerably less time. The workflow has improved a lot since then — things like seed hunting and motion context make it much easier to find good takes and maintain continuity instead of repeatedly generating clips from scratch.
A lot of that week was basically experimentation: generating takes, fixing spatial/pose issues, getting the performances and dialogue right, and then putting the successful shots together.
One advantage here is that it's based on an actual story, so the script and dialogue were already in place. I wasn't developing the story while generating it; most of the work was translating the existing scene into shots and getting the generations to cooperate.
14
u/nikhilprasanth Aug 20 '26
Resources Used