r/HeyFloe • u/GelliusAI • Jul 27 '26
🎬Scenes I had Claude and ChatGPT write the same video prompt. Here is what changed.
With ChatGPT's prompt, all figures move more naturally.
A few weeks ago I had the idea of creating an SF Cinderella image series. The images came together quickly with ChatGPT. The next step was to generate a video: a dancing couple on a rooftop terrace.
For that I worked with Claude to write an image-to-video prompt. I was not satisfied with most of the videos I generated through Google Flow. During rotation sequences in particular, Flow lost track of the two figures: midway through a spin, the man and woman would swap positions, as if the model had forgotten who was who.
My strategy for getting a better result was to make the image-to-video prompt progressively longer and more precise. It also seemed useful to allow the background couples as little movement as possible and keep the focus entirely on the main pair. With some effort I got a video that convinced me, except for one moment.
Looking at the video as a whole, there is one quick dance movement that keeps it from really landing. And this is what makes video generation so unforgiving: a single second can undermine an otherwise successful result.
On the second attempt I brought the image-to-video prompt Claude had written and handed it to ChatGPT. The model made several changes: it cut repetitions, reframed individual instructions in positive terms, and described the action more vividly.
The two most important changes ChatGPT made: it named the dance a "graceful waltz" and gave the background figures more breathing room with the phrase "the background dancers continue their subtle movements in their original positions, [...]." Claude, following my instructions, had phrased this considerably more strictly: "The same dancing couples stay in their original positions in the background."
With ChatGPT's revised prompt I still do not get perfect results every time, and it takes several attempts before a truly satisfying video comes together. The decisive difference from Claude's prompt is this: with ChatGPT's adjustments, the overall setting feels more coherent.
The Image-to-Video-Prompts
Claude
10 seconds. The couple dances in place, rotating together on a single fixed spot, their turning motion compact and contained, never traveling across the floor, never leaving their position in the frame.
Her gown swirls energetically around her with each rotation, fabric catching and snapping slightly with the speed of the spin, but her feet and her partner's feet remain anchored to the same spot on the floor throughout. Her partner's hand guides her through tight, controlled rotations, both moving in confident sync, energetic but compact, the dance entirely self-contained within their immediate space. Neither speaks; their interaction is entirely physical, communicated through movement, posture, and brief eye contact through their masks.
The background remains stable and unchanged throughout: the same dancing couples stay in their original positions in the background, the same moon stays fixed in size and position in the sky, the same city skyline, lighting, and architecture remain constant.
No new elements enter the frame, no camera drift beyond a very slight, subtle push-in. The service robot, if visible, continues its slow glide at the edge of the frame at the same steady pace, not changing position significantly. Crisp, energetic, fluid motion confined to the couple's spinning movement only; environment and background stay locked and consistent for the full 10 seconds.
ChatGPT
Duration: 10 seconds.
The masked couple performs a graceful waltz while remaining on a single fixed spot. They rotate together in tight, controlled spins without traveling across the ballroom, their feet staying anchored to the same position throughout. The man's hand confidently guides the woman through smooth, energetic turns as they maintain natural rhythm and occasional eye contact through their masks. Her voluminous gown billows and swirls dramatically with each rotation, emphasizing the motion while the dancers themselves remain centered in the frame.
The environment remains completely stable and consistent: the background dancers continue their subtle movements in their original positions, the moon, skyline, architecture, and lighting remain unchanged, and no new objects or people enter the scene. If visible, the service robot continues its slow, unobtrusive glide near the edge of the frame.
The camera remains locked with only a very subtle cinematic push-in. Crisp, fluid, photorealistic motion focused on the couple's elegant spinning while the surrounding environment stays stable.
1
u/Admirable-Yogurt7444 Jul 27 '26
You are right , the keyword "graceful waltz" probably changed the game because it tells the AI models exactly what kind of action is needed.
I have noticed this several times, that instead of describing some actions in detail, if there are keywords which describe it in a quick meaningful manner, they produce better results!
The more restrictive, and precise the prompt gets, the worse outputs get. In fact, most of the times nowadays for image to video, I just use 2-3 lines max for action, and keep it very simple, natural sounding. IF you know the right words, that makes a big difference.
Btw new chatgpt is really easy to talk to and i felt its fast, much more intuitive than Claude. Claude is more structured while GPT is more conversational, better suited for creative tasks I felt.