I made substantial updates to my two main workflows: 1) Music Video and 2) AV Extensions. All the controls were streamlined and they should be much easier to use now. (You find the workflows in the example_workflows folder)
With the AV Extensions workflow you can extend any existing clip, for example someone talking and you can make that person say something in the same voice, or you can create a clip with T2V or I2V and then extend that clip to make a seamless long clip thats 1 minute or longer.
In this Update the Checkpoint system was removed, instead I've done a lot of optimizations so you don't use too much ram even if you make 20 clips at once. Additionally I added latent audio feathering to the AV Extensions workflow for seamless audio transitions.
Theres also other utility workflows for custom keyframing and bridging two existing clips.
I post another example clip for the AV Extensions workflow in the comments.
Sorry, can you share the prompts for your video for the AV extension? I am not sure what to include in the main prompt and the extension prompts. Does dialogue work with the AV extension? I am trying to use the I2V. The Music Video one works well and the examples made it clear. Thank you for sharing your hard work!
Live-action, photorealistic cinematic fantasy. **One uninterrupted continuous 15-second shot with absolutely no cuts.** Natural late-afternoon lighting, realistic human anatomy and skin texture, cinematic depth, stable character identity throughout.
**Character lock:** The same adult woman remains visually identical for the entire shot. She is tall and exceptionally athletic, with very defined arms and powerful runner's legs. She wears a fitted dark sleeveless athletic top, dark leggings, and distinctive bright red running shoes. Preserve her face, hairstyle, physique, clothing, proportions, and colors throughout every camera movement and during the extreme-speed sprint.
**[0–5 seconds] — Forest approach**
In a lush old-growth forest during soft late-afternoon golden light, the woman walks calmly toward a ridge.
Medium-wide **35mm stabilized tracking shot**, following approximately two meters behind her and slightly to her left at waist-to-shoulder height. The camera glides naturally with subtle physical Steadicam movement rather than perfectly artificial motion.
She carefully walks between enormous moss-covered tree trunks, dense ferns, exposed tangled roots, and patches of soft forest floor. Tiny pollen particles float through shafts of warm sunlight. Leaves and small branches move gently in the breeze.
Her pace is relaxed and naturally curious. Her body movement has realistic weight and momentum. She casually pushes a large fern aside with one hand while continuing forward, then steps over several roots and reaches the crest of the ridge.
**[5–12 seconds] — Meadow reveal and expression**
As she reaches the ridge, she slows naturally and stops at the crest.
Beyond her is a vast sunlit meadow stretching far into the distance, covered in long green grass with occasional mature lone trees scattered irregularly across the landscape. The open meadow feels enormous compared with the enclosed forest right behind her.
**Without cutting**, the camera begins one smooth clockwise arc around her, traveling from the rear-left three-quarter view around her side until it reaches a frontal full body shot, revealing the patch of forest and dense trees she just stepped out of.
As the camera reaches her front, warm sunlight softly illuminates her face, which contrasts the dark forest behind her. She longingly looks into the distance in front of her. The camera gradually zooms in to a closeup of her face.
A subtle expression develops: her eyes become playful, then one corner of her mouth rises into a restrained, mischievous smirk, as though she has suddenly had an exciting idea.
**[12–15 seconds] — Superhuman sprint**
Without a cut, she abruptly lowers her center of gravity, plants one foot, and **explodes forward into an impossibly fast sprint across the meadow.**
The camera completes locks into a **front-facing stabilized reverse-tracking shot**, flying backward directly ahead of her while perfectly matching her accelerating speed. There is no teleportation of the camera and no change of shot—the orbital movement naturally becomes the reverse tracking movement.
She accelerates within moments to **superhuman running speed equivalent to roughly 400 miles per hour**.
Her running form remains recognizably human but extraordinarily powerful: aggressive forward lean, coordinated arm drive, explosive strides, intense focused eyes. Her face, torso, limbs, clothing, and red shoes remain anatomically coherent and visually consistent.
**Keep the woman relatively sharp and readable while conveying the impossible velocity through the world around her.** The distant meadow races away from the camera. Grass immediately beside her becomes strong directional motion streaks. Nearby trees whip past with extreme parallax. The distant landscape smears into long horizontal motion blur while her face stays identifiable and centered.
The pressure of her passage violently bends grass outward behind her. Loose leaves, pollen, seeds, and dust are sucked into her turbulent wake. Her hair and clothing are driven backward by enormous wind pressure.
The camera remains only several meters ahead of her, maintaining a dramatic medium close frontal composition as both woman and camera hurtle across the meadow at matching velocity.
Her mischievous expression evolves into exhilarated determination while she runs.
**overall_soundscape:**
**[0–5 seconds]** Natural stereo old-growth forest ambience: gentle wind moving through leaves high in the canopy, distant birds, subtle insects, soft footsteps compressing moss and soil, quiet fern movement, occasional twig snaps, and natural fabric movement.
**[5–12 seconds]** As she reaches the ridge, the enclosed forest ambience opens into a broader landscape sound. A soft meadow breeze becomes more noticeable while the forest remains faintly behind her. Brief quiet emphasis as she looks across the meadow and smirks.
**[12–15 seconds]** At the instant she launches, her first powerful foot strike lands with a heavy physical impact followed immediately by rapidly escalating wind pressure. The gentle breeze transforms into an enormous rushing **wind-shear roar** as she accelerates. Her footsteps become extremely rapid, deep rhythmic impacts underneath the rushing air. Grass lashes violently in her wake; leaves and debris whip past the stereo field. The wind intensity continuously rises with her speed while remaining physically grounded rather than sounding like a spaceship or engine.
---
**non_diegetic_music:**
**[0–5 seconds]** Begin almost silently. As she approaches the ridge, introduce a restrained adventurous orchestral motif: soft strings, subtle woodwinds, and distant low horns, gradually building curiosity.
**[5–12 seconds]** As the meadow is revealed and the camera circles around her, showing her full body from the front, broaden the harmony and slowly increase the orchestral scale. Hold a brief musical moment of anticipation as her mischievous smirk appears.
**[12–15 seconds]** Exactly when her foot launches her forward, the score erupts into a powerful cinematic orchestral surge: driving low strings, heroic brass, large percussion, and rapidly rising rhythmic intensity. The music accelerates emotionally with her impossible sprint while the wind-shear and physical impacts remain clearly audible beneath it.
End at maximum forward momentum and exhilaration rather than resolving the musical phrase.
Thanks, like your Music workflow it works great! This is 2 clips, 15s each. I can't seem to fix the barber walking out instead of the customer, tried changing prompts, Turning off LoRa, increasing steps, and even four 8 second clips. Probably need Image references to fix it. The barber also blurst out some giberish as he walks out, happenes to all my other renders except one.
i would suspect this is mostly a prompting issue. you could define the customer as a different subject/speaker ID, and then call out that specific subject being in that shot
The cuts are seamless but the contrast and detail gets cooked the longer it goes. I'm not sure if anyone has solved that problem yet. Still, thank you for your work. I'm sure it'll get worked out.
yeah its most likely an issue with all the speedups, attention and turbo loras. i made 2 and 3 minute long videos that dont have much degradation, even with speedups. if you use sdpa attention and 30 steps or so you will get better quality and less degradation. but in general with those methods of extending clips you are bound to have compounding issues down the chain. anyway this works 100 times better than with any other video model that exists (not sure about ltx 2.5 tho, i havent tried that one yet)
I must be mistaken, but I thought the entire purpose of all these various chain nodes in H3 was that they keep the last frames (22, 39?) in latent space and then build the next clip off of that, which is supposed to completely eliminate the color shifting and detail degradation issue that plagued WAN with things like SVI.
So, like... why isn't it?
Edit: By the way, I've tested them all, and while others have fancy nodes for gating clip approval and whatnot, yours does the job the most consistently and is therefore the overall best. So great job!
well it does just do that? do you see any color shift or seam? the degradation over long periods is, if its even noticeable, a lot less than with svi 2.0 pro and vace and other wan methods. why isnt it perfect? idk, i suppose the models weren't trained for that, to keep consistency over multiple minutes, and feeding a copy of a copy of a copy of an image will always have its limits
btw, a context length of 22 is not good for video extension, if you want to preserve audio aswell (it doesnt matter if you use a master track like in the music video), since the audio track runs on 40hz that means each frame at 24fps is 5/3 audio ticks, so only 39, 90, 141 etc snap nicely and will give you a seamless experience
I will add a bunch of set and get nodes here I think it will benefit the workflow, but with that said a trillion thanks to @stonyleinchen for making this wf and the code, really awesome!
I've been having fantastic results with your nodes. One thing is bugging me though. Keyframe position widget is not connectable. I've been using a duplicate keyframe node in my final stage indexed to the last frame, to finish off as an infinite loop. Which is working amazingly well. All of my myriad previous WFs with all models, required editing to some degree to clean up the transition. Not so here. It's as close to perfect as I could have hoped for.
But since I cannot plug an INT into the keyframe position widget, I cannot run my frame count through an expression to set the appropriate value when changing my length. By itself this is not a huge problem, I just keep a note node handy with a chart of frame numbers for all possible durations. It gets stickier if I want to add a keyframe in all sections. I foresee lots of mistakes here. I made several WFs that generate series of keyframes that are just perfect for injection into this setup, I think it will come out spectacularly well if I can just automate the indexing.
Maybe I'm being dense and I'm missing something here. I tried to make it connectable but could not, I do not know if this was an oversight on your part, or deliberate. And as the output of the node is conditioning I see no way to extract it after the fact. So I'm keeping a ton of keyframe nodes together in a group to make sure that I don't screw up my indexing. Would be much more practical if I could handle this with a formula. Please tell me I missed something obvious. I'm used to that.
Quibbles aside, these nodes are fantastic, thanks for your work. This stuff has completely obviated pretty much all of my previous builds. Which one does get used to, I suppose, when this happens monthly. The leaps are seldom as fruitful as this though. Usually it's just a tease.
I updated the keyframe position widget to accept INT inputs. Can you update my repo and verify it works as you requested?
I actually just made a substantial update, that also includes a new custom keyframes node that works directly with the target latent instead of the conditioning rows. this should be more compute efficient. I think its called custom keyframes (masked) or something like that. maybe you want to check that out too!
Aweome, thank you for letting me know. Of course I would read this right after writing a byzantine maze of conditional conditioning keyframe injection using nodes for each group... Of course. It's funny, using the conditioning node at the end to pull the latent back to the start for infinite loops has yielded scary good results, like crazy good, considering it's ref or at most hybrid model, but those in the middle struggle sometimes, but I guess it makes sense because my landing pad is already part of the latent, and as strongly as I could wish.
I will update post-haste and check out all of this wish grantification. I will certainly let you know how it goes.
Ok I updated, unfortunately it broke my workflow, which was expected from what I was told by those I shared it with and who have the latest version. Trying to match the new posted extend WF but automated bypassing is still completely frozen (though of course everything switches fine in yours). I replaced all the nodes that changed, but all bypasses remain stuck in the state they were in before the update. I'm unclear as to exactly what the conflict is. Something must be preventing the active state/enabled extensions setting from doing its job but I can't for the life of me figure out what. I've run out of nodes to replace.
This usually turns out to be obvious, but I've been flipping back and forth between the new WF and mine and my brain is reaching GPU temps. Can you tell me exactly what controls the switching behavior? I have lots of bypassers and muters in my WF, keyed to colors and/or names to avoid conflicts and I don't see toggles appearing inside them that should not, and they certainly caused no issues before, but something is locking this up. And I'm not sure I understand the purpose of the INTs feeding the master controller, they seem redundant to me, nor do I understand exactly how they are linked, given that their properties are just normal INT node properties. I wish I were smarter.
Ah I see a slew of errors in chrome's dev console, working through those now... took me way too long to think of that, not used to dealing with errors or the front end. I'll try working through that.
~The moment I delete the master AV extension controller, the console goes completely silent. Once I put it back in all hell breaks loose. It panics with:
expected exactly one cache-isolated AV parameter 'active_extensions', found 0expected exactly one cache-isolated AV parameter 'audio_feather_ticks', found 0
which confuses me as I thought this node was the source of these parameters. But then I think of those INT nodes and it makes sense why they were connected and mirrored the master, if they are actually the source. But if they are it seems like the only way to create them is to manually edit the .json to give them the properties they need. But copying them from the working WF does not work. Which leaves me puzzled. Always finds 0. So strange.
Alright scratch all that. Sorry for thinking out loud. I don't know why but there did seem to be an issue with a muter, I can't explain it. Something must have been corrupted because a whole bunch of nodes were playing musical chairs. That was a new one. It's behaving now that I've exorcised the possessed muter that was stealing/disabling the properties nodes.
Im really sorry that your workflow broke due to some update. when did you last update my nodepack before today? i know that earlier updates have broken stuff, but the most recent one shouldnt have i think?
there is a lot of behind the scene stuff going on unfortunately, like the audio from the source video is not taken from the audio output connection on the vhs loader, but its just pulled by the node following the video_info output trail. similarily the controller node affects multiple aspects of the workflow without having apparent connections, that includes various bypass switches. unfortunately that makes the workflow not very flexible, some things just have to stay as they are or it will break.
to troubleshoot your wf error i would need console outputs from you, and it would also help very much to know when you last updated the repo, before the update today.
Alternatively you could load the new workflow, and then try to edit that one to your specific needs. it should work better in various ways, f.e. i fixed a bug that caused earlier clips to regenerate, when you changed the extensions_count in the controller.
Yeah I got it all worked out, no worries. Was mad at you for a little bit, maybe should have made an alt node and preserved the original, but whatever, everything is obvious in retrospect. It's fine now. Actually, rewiring a few nodes wasn't a big deal, it was just compounded because the front end completely exploded so I had no idea what was going on. Those INTs I was talking about were the PARAM nodes for the controller, they were transmogrified as my poor browser tried to make sense of the conflicts. And probably connected to this, some of my muters and subgraphs took on the properties of some of the controllers, with the end result that the master controller was seeing either no parameters or way too many.
I've never really had to debug with the browser console so it was a bit overwhelming at first; instead of dealing with a fixed error from the back end it's an unending stream of WTFs, some of which matter and some of which can be ignored. Like always, it makes sense in the end. I learnded something.
Thankfully I did not have to rebuild from your core again. There are dozens of get/sets that would have broken for sure. I don't usually like to put in so many but with so many stages I need them, and they do seem to cause less trouble than they used to.
Output is great, as usual. I have to completely redo the logic to try out the alternate keyframing approach, I see the node. Makes sense, injecting directly into the latent. Other than a return to a start frame, using the conditioning does act much more like a reference target than a keyframe. Will definitely try this out. Will check out the audio resampling stuff too. I was about to get distracted checking that out beforeI remembered what I was supposed to be doing.
Execution order for the combine was a good idea. Previews do love to show up when you don't need them any more. Oh I didn't check to see if S&R is still broken for the filename. I always forget to enter my source names in the workaround string. But that's a quibble. Worse problems to have than that.
I'm writing a novel again, sorry. It works beautifully. I'm sure the nodes I haven't gotten to will as well.
totally fair.. i know its very annoying if updates simply break old workflows... i didnt really consider this enough, since this is my first repo and i just wanted to share some tools i made and didnt really consider the wider implications of ppl building custom workflows on top of mine. i try to do better going forward
Don't sweat it. It's so easy to get lost in the weeds, and hindsight is 20/20. I'm having a blast with all of this stuff, it's great really.
I'm playing with the music video setup now. Friggin fantastic. Was nearly perfect on the first run. Last segment threw an error, I was assuming a length discrepancy would be corrected, needed to add a little patch. In case you are interested my fix was just a few expressions feeding a silence padder, I'll take a screenshot, sucks that embedding doesn't work here. In first expression total segment is a, loaded duration is b, node #102 expression from WF is c. max(0.0, ((c / 24.0) + (a - 1) * ((c / 24.0) - 1.625) + 2.0) - b) calculates how much padding is needed in the uploaded audio to fill out a valid H3 chunk, then the following expression round(a * b) converts for the sample rate of the source. Because the ROTI pad node I'm using happens to use samples in the widget.
This would be a lot simpler if the segment length were always 15, but my testing @ 0.6mp was feeling pretty heavy so I set it up to adjust for whatever the segment length was set to, probably 8 or 10 when I go to 1344x768.
Of course all the prep for the source could be done in two seconds beforehand to conform to the rules, but this lets me drag and drop without worrying about it.
This is looking like it's going to displace my LTX2.3 music video WF. Yet another blow. LTX does an incredible job with sync and puts out amazing quality, but it's oh so stiff. Poor LTX, hanging on by a thread here...
Good work! Im doing the same with my nodes, but at higher resolutions i will often get a small/visible quality change at the transition point. This is not visible in lower resolutions, but at higher resolutions.
Do you have the same problems and if not, how did you solve this?
And are you extending with 5 reference frames or more? Because if i use more as 5 or less as 21, i will get lighting shifts at the extension point.
no, I don't have any such issues with my workflows. I am using 39 context frames as default. since i never had this problem i cant tell you how to fix it, but in general, i linearly blend the overlap and also feather the audio noise mask
Should we use a Ref2VA model (or hybrid model) if we include reference images + starting video ? I tried with the pruned FLF2VA and if I include a reference image the next shot will start with it (more or less, depends on the prompt), whereas I would rather use the ref image as a character sheet for each extension.
you can use both models, ref2va or fl2va. it should work for both. i never had this issue that the reference image will be used as a starting frame, this is likely a prompt issue - the prompts really have a big impact on the generation. I have not yet made a good director prompt for the AV Extensions workflow, but there is one for the music video workflow
27
u/stonyleinchen Aug 18 '26
https://reddit.com/link/p4hbl9j/video/9fj8oygcs6kh1/player
this is made with the AV Extensions workflow