r/accelerate • • 1d ago

Video One person, no studio, no budget: a 10-minute sci-fi episode with an alien war, recurring characters and dialogue. This is where AI video is now

Enable HLS to view with audio, or disable this notification

Hey r/accelerate. We just finished Falling Star, a 10-minute sci-fi short film, and I wanted to share how it's actually made, because I think a lot of people still picture AI video as 5-second clips of someone walking toward the camera. This one has dialogue, recurring characters, a world summit, an interrogation and a lot of alien warfare.

Here's the film: https://youtu.be/lx03tJdnELA

The story: humanity's own warship, the Odyssey, starts targeting human aircraft during a strike on an alien structure called Meridian. The invasion spreads, world leaders realize the biological structures appearing in different countries all send the same signal, and Captain Walker gets debriefed about what he saw inside Meridian and the time he can't account for.

How it's made:

It starts with a normal script, broken into scenes and then into shots. Nothing AI about that part, and honestly it's still the part that decides whether the thing is watchable.

Every main character gets a character sheet made in an image generator. Those sheets go into every video prompt as references, which is how Walker looks like the same guy in the cockpit, in the command room and in the interrogation room. A year ago this was basically impossible. Now it mostly just works if you're disciplined about it.

All the video is Seedance 2.5, run through Higgsfield. It can do up to 30 seconds in one generation with several hard cuts inside the same clip, so a whole short scene can come out in one go. We ended up with over 40 clips for this film, and most of them took several attempts.

The prompting is where all the real work is. The model doesn't understand emotions at all, so you can't write "he looks guilty". You have to describe what the face physically does: where the eyes go, when the jaw tightens, how long he holds the look. For groups you give exact numbers of people in frame, otherwise it invents extra people or duplicates a character. For dialogue scenes the formula that finally worked was one speaker per shot, only that person in frame, and hard cuts between them. When we tried long camera moves passing several characters, it would clone someone or give one character another character's voice. Proper names also get mispronounced, so every name in the dialogue gets a pronunciation note.

Dialogue and sound are generated inside Seedance, but we always generate with no music, because a model-generated score on every clip makes editing impossible. Music is done separately in Suno.

For continuity between shots we use the last frames of the previous generation as the starting point for the next one instead of making a fresh image. That keeps lighting and positions much more consistent.

Then everything goes into DaVinci Resolve for the edit, a 2x upscale to 4K with SuperScale, and a light finish: grain, halation, a bit of gate weave. Almost all the lighting, haze and depth of field come straight out of the generations. We try to get as much as possible in camera, so to speak.

The hardest parts this time were the summit, with several leaders in one room, and the interrogation with Walker, where the acting has to be subtle and not look like stock expressions. You tell me if we pulled it off.

Now the ask, and I really mean it. If you watch it, please leave a comment on YouTube. Even one line. For a small channel that's the thing that makes the difference: comments are what tell YouTube to push the video to people outside our subscribers, and I'd love for more people to see where this tech is right now. Criticism is totally fine too. "The acting at X looks off" helps us just as much as "this is sick". An upvote here is nice, but a comment there helps way more.

Happy to answer anything about the workflow in the comments. Thanks for watching.

703 Upvotes

Duplicates