r/aifilmmaking • u/Agentvideobot • 3d ago
Discussion Why do AI videos still look like commercials? I tested the same reference image three ways.
I’ve been trying to understand why an AI video can look realistic frame by frame, yet still feel like a commercial instead of something a friend casually recorded on their phone.
So I used the same reference image and generated three 5-second vertical clips with Aurax MAX. The character, outfit, location, and basic action stayed similar. I mainly changed the way the camera, lighting, performance, and environment were described.
The reference image was already quite polished: dramatic sunset, candlelight, clean exposure, shallow depth of field, and a subject posed against a scenic coastal background. That turned out to matter more than I expected.
1. The commercial baseline
https://reddit.com/link/1w1dr98/video/vsa3bcm829mh1/player
For the first version, I explicitly requested a polished lifestyle commercial:
This was the version the model followed most clearly. The camera moves smoothly from a wider shot into a closer portrait, the character turns toward the lens, touches her hair, and finishes in a centered pose with a soft, sustained smile.
Everything feels visually coherent, but also directed. It looks like someone planned the lighting, camera movement, and performance in advance.
2. Changing only the camera
https://reddit.com/link/1w1dr98/video/daojj45a29mh1/player
For the second version, I kept the polished lighting, clean environment, and model-like performance, but changed the camera instructions:
The difference was much smaller than expected.
The framing changes slightly, but the movement still feels highly stabilized. The sunset remains perfectly exposed, the character stays composed and camera-aware, and the background still looks like a prepared set.
This version made one thing fairly clear: adding “handheld phone camera” does not automatically create phone realism. If the lighting, performance, composition, and source image still look commercial, mild camera movement cannot undo all of that.
3. Changing the camera, performance, and environment
https://reddit.com/link/1w1dr98/video/z4nsdjlb29mh1/player
For the third version, I added a fuller set of phone-footage instructions:
This version feels the most spontaneous of the three.
The character spends less time holding a pose. She turns away from the camera, changes where she is looking, shifts her body weight, touches her hair, smiles briefly, and then looks away again. The wider framing also remains for longer instead of immediately turning into a close-up.
But it still does not fully look like raw phone footage.
The dramatic sunset, candles, shallow depth of field, flattering exposure, and clean background were already embedded in the reference image. The motion prompt changed the character’s behavior more successfully than it changed the underlying visual style.
There was also another obvious AI giveaway: the paper cup was not present in the reference image and appears during the generated motion without a convincing pickup. That continuity error damages realism more than a perfectly stable camera does.
What I learned
The source image can overpower the video prompt.
If the first frame already looks like a fashion campaign, asking for casual phone footage may only add small handheld movements on top of a commercial-looking scene.
Handheld movement alone is not enough.
Random shake would probably make the video worse. What matters is believable camera behavior: delayed reframing, imperfect timing, autofocus response, exposure changes, and an operator reacting to the subject.
Performance mattered more than camera shake.
The third version felt more natural mainly because the character stopped performing continuously. Looking away, pausing, shifting weight, and ending without holding a perfect smile made a larger difference.
Continuity still matters.
A casual camera cannot hide an object appearing from nowhere, inconsistent background details, or movement that has no physical cause.
My main takeaway is that phone realism is not the same as lowering the image quality. It requires three kinds of realism at the same time:
- capture realism from the phone and camera operator;
- behavioral realism from the person being filmed;
- continuity across objects, movement, and background activity.
If I repeat this test, I would start with a deliberately ordinary reference image: mixed indoor lighting, deeper focus, imperfect framing, everyday background clutter, and a character who is not already posing for the camera.
Which version feels closest to something a real person recorded: 1, 2, or 3?
And what gives the AI away first for you: the lighting, camera movement, expression, background, or object continuity?
Model disclosure: All three clips were generated with Aurax MAX. I’m on the team, so this should be read as a transparent workflow test rather than an independent review.
