← Back to blog
AiAbout 7 min read

The Two-Frame Method: How First and Last Frame Interpolation Changed AI Video Planning

Published Oct 10, 2026
The Two-Frame Method: How First and Last Frame Interpolation Changed AI Video Planning

Google updated Veo 3.1's First and Last Frame to Video feature on October 9. You upload two stills, the opening frame and the closing frame, and the model generates the motion in between. Output goes up to 4K at 60fps with native audio, it supports 9:16 vertical, and it accepts three or four reference images to keep characters and style consistent.

On its own that sounds like a small addition. In practice it changes how people plan a shot, because it puts the two decisions that actually matter back at the front of the process.

The problem with prompt-only video

For most of the past two years, AI video worked like a slot machine. You wrote a paragraph, waited, and watched something come out. The result was often good enough to post and rarely good enough to place. If the character drifted or the camera did something strange, you rewrote the prompt and tried again.

Creators learned to work around this. They generated a strong still, then asked a model to animate it, which is image-to-video. That fixed the opening frame, but nobody controlled where the shot ended. Clips tended to either run out of energy or invent an ending nobody asked for.

Defining both ends removes a chunk of that guesswork. You decide where the audience starts looking and where they finish looking. The model's job is narrower: connect the two.

Why storyboard thinking came back

Animation studios have worked this way for a century. An animator draws a key pose, then another key pose, and fills the drawings between them. The craft is in choosing the poses, not in drawing every frame. AI video is arriving at the same division of labor by a different route.

This is not unique to Veo. The controllability race has been building for months. Kling 4.0 shipped with ten keyframes and 30-second native clips. Seedance 2.5 can re-shoot existing footage from a new camera angle. Wan 3.0 topped Artificial Analysis's AA-Video-T2V v2.0 benchmark with an Elo of 1157 at roughly $12 per minute. Vidu's Q4 preview, released October 7, starts at about $0.014 per second and takes up to 15 image references. What the strongest models share is less about raw quality and more about how precisely you can steer them.

A timeline with two keyframes connected by a glowing motion arc

What a working setup looks like

The two-frame method only pays off if the two frames are worth connecting. A few habits help.

Generate the frames with an image model instead of a screenshot from a previous clip. You want control over composition, lighting and wardrobe on both ends. Google's Veo accepts reference images specifically so a face or a product survives the transition, and that only works if the references are consistent to begin with.

Keep the camera move simple. Interpolation handles a push in, a pull back, a reveal, or a slow drift well. It struggles when you ask for a hard cut inside a single clip, because there is no cut to interpolate. If a shot needs a cut, make two clips.

Match the audio plan to the shot. Veo 3.1 synthesizes native audio, which is useful for ambient sound and simple effects. If the scene needs a specific line of dialogue or a licensed track, plan to add it afterward rather than fight the generated sound.

Reserve the method for shots that carry a story beat. There is no reason to spend two well-built frames on a two-second transition. Use it where the audience is meant to notice the change.

What two extra frames actually cost

Planning in frames is a decision with a price tag, and the price is easier to read now than it was a year ago. The video market has been racing downward on cost per second. Vidu's Q4 preview, released on October 7, starts at about $0.014 per second and takes up to fifteen image references. xAI's Grok Imagine Video 1.5 Lite, added to its API on October 8, lists $0.02 per second at 480p and rises to $0.14 at 1080p. HeyGen's new general video model went out at an October promotional rate of one cent per second. Wan 3.0, which took the top spot on the Artificial Analysis video benchmark at an Elo of 1157, is listed around $12 per minute, which works out to twenty cents a second.

Two things follow. First, the generation itself is rarely the expensive part now. A few seconds of interpolated motion costs cents, not dollars. Second, the scarce resource is the frames. Generating and selecting two strong stills takes a person's time, and that time does not get cheaper when the API gets cheaper. The economics reward planning, not volume, which is the opposite of how prompt-only tools trained people to behave.

A workflow that holds up on a real job

Here is how the method fits a small production without changing everything. Start with the concept and decide the single beat the shot has to land. Build the opening frame and the closing frame as images, using the same references so faces, wardrobe and props match. Generate the interpolation, then watch it once for structure and once for details. If the motion curves wrong, fix the frames rather than the prompt, because the frames are what the model is actually reading. Add or replace audio depending on the shot. Then cut it into the sequence and judge it in context, because a clip that looks impressive alone can still feel slow next to its neighbors.

The discipline here is that every correction is a decision about the story, not a roll of the dice. That is the part people miss when they first try the feature. It feels like a shortcut, but it quietly reinstates the pre-production that the first generation of AI video tools made people skip.

Where a normal camera still wins

None of this argues that shooting is obsolete. Interpolation is strongest when the two ends are strong and the motion between them is simple, which covers a lot of product shots, transitions, title cards and atmospheric establishing frames. It is weakest when reality is doing something specific and unpredictable: a real horse at full gallop, a crowd reacting, food being cooked in a way that has to read as true. For those, a camera plus a competent editor still ends the argument.

The sensible position is boring. Use interpolation where control matters and the motion is knowable, and use a camera where the scene needs to be real. Most working creators end up with both in the same project, and that mix is not a compromise. It is just production.

The honest limits

Interpolation is still interpolation. If the two frames imply a physically complicated in-between, like a person turning a full circle or a liquid changing shape, the model will invent something and it will not always be right. Keep the implied motion within what a camera could plausibly record.

It also pushes cost up. Two polished stills plus a generated clip is more work than one prompt. At the per-second prices now on the market, that is more affordable than it was a year ago, but the planning time is real. The method rewards people who already think in shots.

And there is still no guarantee on the middle. What the feature really sells is a smaller search space, not a solved problem. That is a meaningful step. It means the difference between a usable clip and a memorable one is now mostly about the frames you choose, which is a decision a person makes.

What to watch next

Expect the same pattern to spread. Once one major model treats frames as the primary input, competitors tend to follow, and the interesting question becomes whose interpolation looks most natural at the seams. Watch for tools that let you add a middle keyframe, which would give directors a third anchor without leaving the timeline.

For now, the practical takeaway is unglamorous. If you make AI video, stop treating the prompt as the main creative act. Build the frames first, pick the two that tell the story, and let the model handle the trip between them.

Related articles