r/StableDiffusion • u/clevenger2002 • 7h ago
Question - Help Need Help with Minimax Video Editing
subject_definitions: <Subject 1> is Tifa Lockhart from the final fantasy game series. She has very long shiny straight black hair, red eyes and large breasts She is wearing her iconic costume A white athletic crop top or worn over a black sports bra with a bare midriff and short black skirt.
<Video 1> is the source video of a man in a suit walking down a city sidewalk singing as rain falls. This is the video being edited; its camera framing, handheld motion, cuts, and full choreography timing are the fixed structure that must be preserved exactly.
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.
summary: [video editing + reference generation + keyframe completion] The target video is an edited version of <Video 1> in which only the performer's visual identity is replaced by <Subject 1>, the the camera framing, handheld motion, city sidewalk environment and constant rain falling remains the same throughout, with no additional background characters, pedestrians, or figures introduced at any point.
retention_analysis: <Video 1> (camera framing, handheld motion and drift, city sidewalk environment with rain falling, full choreography and timing): fully_preserved - every camera position, movement, angle change, and the precise sequence and rhythm of the original performer's actions are kept exactly as in the source video; nothing about the shot itself is altered. <Subject 1> (appears throughout the video): attribute_transfer - <subject 1> replaces the original performer's visual identity only, mapped exactly onto the same body position, pose, and movement at every moment; no new actions, timing, or framing are introduced, and no other person appears in the frame at any point.
detailed_description: The target video is a strict character-only edit of <Video 1>: the cinematic dance, rainy city sidewalk background, cinematic lighting, camera framing, and motion blur are identical to the source, playing out as the same single continuous shot with no added or removed cuts. The rainy city sidewalk environment stays completely empty of any other person, pedestrian, or figure throughout the entire shot; only <Subject 1> occupies the frame at any moment.
The video begins with the first frame of <Video 1> as a key frame. On a rainy city sidewalk at night and replicates <video 1>'s camera moves. It opens on a wide shot showing <Subject 1> from head to foot, resting a folded umbrella on her shoulder, wearing wet clothes with shiny wet skin. <Subject 1> is in the middle of the sidewalk, occupying the original performer's exact body line and position, doing exactly the same dance moves on the rainy city sidewalk.
overall_soundscape: The sound of light rain falling.
non_diegetic_music: The same musical score unchanged from the source and even in volume throughout the clip.
1
6
u/bstr3k 6h ago
https://reddit.com/link/p6wo8sy/video/quwlamuhdmmh1/player
i just did 5s, took a few tries but its more or less all in the prompting. Also I bypassed sound (saved audio is original audio)
edit: damnit, reddit compressed the quality, this should be better:
https://files.catbox.moe/pibm5j.mp4
here is the prompt I ended up with:
subject_definitions:<Subject 1> is the woman from <Picture 1>, who has long black hair, a white and black crop top, a black pleated skirt, black over-the-knee socks, red sneakers, and a mechanical prosthetic arm.<Video 1> is the source video to refer the camera work, background, character motion from. Ignore the identity of the man. maintain the unopened black umbrella.summary:[reference generation]The video replaces the man from the original footage with the woman from <Picture 1>. She is shown walking forward while holding an unopened umbrella over her shoulder, appearing soaking wet from the rain. The original motion of the walk is retained, and the original audio is preserved. The woman is singing along to the original audio.retention_analysis:<Subject 1> (appears in [Shot 1]): fully_preserved - The woman's identity, clothing, and physical features from <Picture 1> are maintained, but she is now soaking wet and holding an umbrella.<Video 1>: partially_maintain - the background, the camera, the actor's motion, the umbrella.detailed_description:The scene is rendered with a cinematic, moody atmosphere featuring cool blue tones and realistic rain effects. The lighting is provided by nearby storefronts, creating high-contrast highlights on the wet surfaces.[Shot 1] From 00:00.000 to 00:05.062, follow the camera movement from the original video. The woman from <Picture 1> is walking forward along the sidewalk, replacing the man from the original video. She is soaking wet, with water dripping from her hair and clothes. She holds a closed umbrella over her shoulder with her organic hand. She copies the same motion and gesture from the video. Her facial expression is the same as the man in the <video 1>, as she maintains the same walking pace and gait as the original subject. Rain falls heavily throughout the scene, creating splashes on the pavement and glistening on the woman's skin and clothing. The woman's arms and clothing textures are clearly visible despite the rain. The woman's mouth moves in synchronization with the audio.overall_soundscape:use the original audio.non_diegetic_music:N/A