EditorNodesPricingBlog

START AND END FRAME

Fix both ends, and generate the middle

Supply the first frame and the last frame, and the model produces the shot that connects them. Set how long the transition takes, how much movement it carries, and whether it comes back with sound.

Updated:

What start and end frame video does

Start-End Frame Video takes two images and generates the clip that travels from one to the other. Both ends are inputs rather than outcomes, so the only thing the model decides is the route.

That constraint is the point. Text to video decides everything and image to video decides the ending. Here the model is given the least freedom of the three generation types, which makes it the most predictable when you already know what the shot has to do.

The start and end frame form, control by control

Nine controls, and two of them, movement amplitude and the pair of audio toggles, do not exist in the other video forms.

Generation type dropdown in the video workspace with Start-End Frame Video selected

Start-End Frame Video is the third entry in the generation type list.

1

Pick the generation type

Start-End Frame Video sits between Text to Video and Draw to Video in the type dropdown. Selecting it rebuilds the form with two required image blocks instead of one.

That second required input is the whole distinction. Image to Video treats a last frame as optional, and here both ends of the shot are inputs the generation cannot run without.

2

Choose between two models

The roster is short. Vidu Q3 Pro is the default at 540p to 1080p and 1 to 16 seconds, and Seedance 2 is the alternative at 720p to 1080p and 4 to 15 seconds.

Vidu Q3 Pro is the more flexible of the two on length, reaching both shorter and longer than Seedance 2. Where a transition needs to be very brief, a one second Vidu generation is available and Seedance will not go below four.

Start-end frame model dropdown showing Vidu Q3 Pro and Seedance 2 with their limits

Two models here, against thirteen in text to video.

Start Frame Image Input block offering upload, previous task, AI actor, AI character and URL

The start frame accepts files up to 300 MB.

3

Upload the start frame

The start frame block takes an upload, a previous completed task, a saved AI Actor, a saved AI Character or a direct image URL. It accepts JPEG, PNG, WEBP, GIF and AVIF at a minimum of 300 by 300 pixels.

Its upload ceiling is 300 MB, which is three times what the end frame block allows. If you are working from large source files, the asymmetry is worth knowing before you hit it.

4

Upload the end frame

The end frame block mirrors the start frame block, with the same five input methods and the same accepted formats. Its upload limit is 100 MB rather than 300 MB.

The two images need to be plausibly the same scene for the transition to read as a shot rather than a dissolve. Frames that share a subject, a lighting setup and an approximate camera position give the model a path to interpolate along.

End Frame Image Input block with the same five input methods as the start frame

The end frame takes the same five inputs, with a 100 MB ceiling.

Detailed Prompt field in the start-end frame form with a 2500 character counter

The prompt guides the route between the two frames rather than the frames themselves.

5

Describe the path, not the endpoints

Both ends are already fixed by your images, so a prompt describing what they contain is spent on information the model already has. The useful content is what happens in between.

Naming the route matters when more than one is plausible. Two frames of the same room from different angles could be a pan, a cut on movement or a dolly, and saying which one you want is the difference between three attempts and one.

6

Set how long the transition takes

The slider runs from 1 to 16 seconds on Vidu Q3 Pro and starts at 5. Duration has an unusually direct effect here, because it sets the speed of the transition rather than just the length of the clip.

The same two frames over 2 seconds and over 12 seconds are different shots. A short duration reads as a snap or a whip, and a long one as a slow reveal, so this control is a creative decision rather than a technical one.

Video Duration slider in the start-end frame form running from 1 to 16 seconds

Duration from 1 to 16 seconds on Vidu Q3 Pro, defaulting to 5.

Movement Amplitude control with Auto, Small, Medium and Large options

Movement amplitude, a control that exists here and not in image to video.

7

Set the movement amplitude

Movement Amplitude offers Auto, Small, Medium and Large. It governs how much motion the model introduces on the way between your two frames, separately from how far apart those frames are.

Auto is a reasonable default. Small is the setting to reach for when the two frames are already close and the automatic choice is adding movement the shot does not need, and Large suits a transition that should feel energetic rather than smooth.

8

Set the output resolution

Vidu Q3 Pro offers 540p, 720p and 1080p, defaulting to 720p. Seedance 2 starts at 720p and does not offer the 540p option.

540p exists to make testing cheap. Whether two frames interpolate into a coherent shot is visible at 540p, so establishing that first and regenerating at 1080p is the sensible order.

Resolution control offering 540p, 720p and 1080p with 720p selected

Three resolutions on Vidu Q3 Pro, with 720p selected by default.

Generate Audio and Background Music toggles in the start-end frame form

Two independent audio toggles, both on by default on this model.

9

Decide about sound

Two separate toggles control audio. Generate Audio produces sound tied to what is happening in the picture, and Background Music adds a score underneath it.

For a clip heading into an edit with its own sound design, turning both off avoids generating audio you will discard. For a standalone piece the generated track saves a step.

Models available for start and end frame

Two models rather than the full video roster. The interpolation task is narrower than open generation, and the models offered here are the ones built for it.

ModelBest forNotes
Vidu Q3 Pro
Default
Most transitionsThe default. 540p to 1080p, 1 to 16 seconds. The only option for very short or very long transitions, and the only one offering 540p for cheap tests.
Seedance 2An alternative read on the same frames720p to 1080p, 4 to 15 seconds. Worth trying when Vidu interprets the path between two frames in a way you did not intend.

Resolution and duration figures come from the model picker in the workspace, which prints them beside each model name. The controls in the form change with the model: 540p is offered on Vidu Q3 Pro and not on Seedance 2, and the duration slider takes its end points from whichever model is selected.

Settings reference

Every control in the start and end frame form, what it changes and the values it accepts.

Start Frame Image Input

The first frame, required. Accepts an upload, a previous completed task, a saved AI Actor, a saved AI Character or a direct image URL.

JPEG, PNG, WEBP, GIF, AVIF up to 300 MB, minimum 300 by 300 pixels

End Frame Image Input

The final frame, also required, with the same five input methods. Its upload ceiling is lower than the start frame block.

Same formats, up to 100 MB

Detailed Prompt

Guides the route between the two frames. Describing the contents of either frame is wasted, since the model already has both images.

Up to 2,500 characters

Model

Two options rather than the full video roster. Vidu Q3 Pro spans a wider duration range and adds a 540p output option.

Vidu Q3 Pro or Seedance 2

Video Duration

How long the transition takes, which sets its speed rather than only its length. The same pair of frames reads very differently at 2 seconds and at 12.

1 to 16 seconds on Vidu Q3 Pro, 4 to 15 on Seedance 2

Movement Amplitude

How much motion the model introduces on the way between the frames, independent of how far apart they are. Absent from the image to video form.

Auto, Small, Medium, Large

Resolution

Output resolution. 540p is the cheap test setting, since whether two frames interpolate coherently is already visible there.

540p, 720p, 1080p on Vidu Q3 Pro

Generate Audio and Background Music

Two independent toggles. One produces sound tied to the picture, the other adds a score underneath it.

Two toggles, on by default

Choosing two frames that interpolate

The quality of the result is decided mostly by the pair of images rather than by the settings. Frames that could plausibly be connected by a single continuous shot give the model a path to follow. Frames that could not leave it inventing one, and the output reads as a dissolve between two pictures.

Three things carry most of that plausibility. A shared subject, so there is something to track between the frames. A consistent lighting setup, so the model is not asked to change the time of day mid shot. And a camera position that has moved rather than teleported, so the change is a travel the model can render.

Generating both frames in the image workspace with the same seed and prompt, changing only the element that should move, is the reliable way to produce a pair like that. Both frames can then be pulled in through Previous Task without leaving the Studio.

Duration as a creative control

In text to video, duration sets how much clip you get. Here it sets how fast the transition happens, because the distance between your two frames is fixed and the duration decides how long the model has to cover it.

The same pair of frames at 2 seconds reads as a whip or a snap. At 12 seconds it reads as a slow reveal, and the model has room to add detail in the middle that a fast version never shows. Neither is more correct, but they are different shots and the control is where you choose between them.

Movement amplitude interacts with this. A long duration at Large amplitude gives the model both time and licence to wander, which is right for a dreamlike transition and wrong for a clean product reveal. Pairing a long duration with Small amplitude keeps the shot slow without letting it drift.

Model tutorials

Deeper coverage of the models available here and of the wider video roster.

Start and end frame questions

What is start and end frame video for?

Shots where you know both what the audience sees first and what they see last, and the only open question is the movement between them. Reveals, morphs, before and after comparisons, and transitions that have to hand off cleanly to the next clip in an edit are the usual cases.

Similar enough that a single continuous shot could plausibly connect them. Sharing a subject, a lighting setup and an approximate camera position gives the model a path to interpolate along. Two unrelated images will still produce a clip, but it reads as a dissolve between two pictures rather than as a shot.

The start frame block accepts files up to 300 MB and the end frame block up to 100 MB. Both take JPEG, PNG, WEBP, GIF and AVIF at a minimum of 300 by 300 pixels, so the difference only matters when you are working from unusually large source files.

How much motion the model introduces on the way between your two frames, separately from how far apart those frames actually are. Auto is a sensible default. Small suppresses movement the shot does not need when the frames are already close together, and Large produces a more energetic transition.

Image to Video treats the last frame as an optional constraint on an animation, and it does not offer a movement amplitude control. Start-End Frame Video requires both images and is built around the transition itself. Where both end points matter more than the path, this is the right type.

Duration sets the speed of the transition here rather than just the length of the clip, so it is a creative decision. Two seconds between the same pair of frames reads as a snap, and twelve reads as a slow reveal. Vidu Q3 Pro is the only model that goes below four seconds.

Two independent toggles control this, one for sound tied to the picture and one for background music. If the clip is heading into an edit with its own sound design, turn both off rather than generating audio you will discard.

Video models charge per second of output and higher resolutions cost more, so the figure moves with the model, the duration and the resolution. The exact cost is shown in the form before you submit. Failed tasks are refunded automatically during credit reconciliation.

Other ways to make a video

Start-End Frame Video is one of six routes into the video workspace.