EditorNodesPricingBlog

IMAGE TO VIDEO

Start from a frame, then move it

Give the model a still and describe only the movement. Thirteen models, an optional last frame to pin where the shot ends, durations up to 30 seconds, and output to 4K. Any image you already generated can be picked without downloading it first.

Updated:

What image to video does

Image to Video animates a still you supply, optionally guided by a prompt. The first frame is fixed before generation starts, so composition, wardrobe, lighting and framing are settled and the model is left with one job. Deciding how the shot moves.

That is the practical argument for generating a still first and animating it second, even when text to video could have produced the same shot in one pass. A still is cheaper to reject than a clip, and it gives you an approval step in the middle of the process rather than at the end of it.

The image to video form, control by control

Seven controls, and the two image blocks at the top account for most of what the result looks like.

Generation type dropdown in the video workspace with Image to Video selected

Image to Video sits at the top of the generation type list.

1

Pick the generation type

Image to Video is the first entry in the type dropdown, above Text to Video, Start-End Frame Video, Draw to Video and Video Enhancer. It is the type to reach for whenever the first frame already exists.

Selecting it rebuilds the form with an image source block at the top, which is the structural difference from Text to Video. Everything below that block behaves the same way.

2

Choose a model, and read its limits

The image to video list adds the Vidu Q3 Drama and Vidu Q3 Turbo variants that text to video does not offer. Vidu Q3 Drama reads 720p to 1080p and 2s to 30s, which is the longest single generation available anywhere in the video workspace.

Some models behave differently here than they do in text to video. Seedance 2.0 Mini reads 480p to 4K in this list against 480p to 1080p in the text to video one, so the same model name is not the same ceiling across generation types.

Image to video model dropdown listing each model with resolution ceiling and duration range

The image to video roster is longer than the text to video one, and the limits differ per model.

Image Source block offering upload, previous task, AI actor, AI character and image URL

Five ways to supply the first frame. Previous Task is what makes chaining generations practical.

3

Supply the first frame

The Image Source block takes an upload, a previous task from your own history, a saved AI Actor, a saved AI Character, or a direct image URL. The drop zone accepts JPEG, PNG, WEBP, GIF and AVIF up to 100 MB, with a minimum of 300 by 300 pixels.

Previous Task is the option worth knowing about. It lets a still generated in the image workspace become the first frame here without a download and re-upload in between, which is how a still and a clip end up part of one sequence rather than two separate jobs.

4

Add a last frame if you have one

A second, optional block accepts a last frame image using the same five input methods. Leaving it empty lets the model decide where the clip lands. Filling it pins the final frame.

This is worth distinguishing from Start-End Frame Video, which is a separate generation type built around the same idea. Here the last frame is an optional constraint on an animation, and there the transition between two images is the entire job.

Optional Last Frame Image block with the same five input methods as the first frame

The optional last frame. Supplying it constrains where the shot ends up.

Detailed Prompt field in the image to video form with a character counter

The prompt is optional here, and it should describe motion rather than the scene.

5

Describe the movement only

The prompt is optional in image to video, because the model already has the scene. When you do write one, everything spent describing what is visible in the frame is wasted, since the frame is already settled.

Use the space on movement instead. Which elements move, in what direction, at what speed, and what the camera does over the clip. A short prompt naming one clear camera behaviour usually beats a long one restating the picture.

6

Set the duration

The slider runs between the end points of the selected model, which the picker printed beside its name. Vidu Q3 Pro accepts as little as 1 second, and Vidu Q3 Drama reaches 30.

Animation faults show up early. A movement that looks wrong in the first two seconds will still look wrong at fifteen, so testing at the short end is the cheap way to find out whether the prompt and the frame agree.

Duration slider in the image to video form running from 5 to 15 seconds

Duration, with end points set by the selected model.

Video resolution control in the image to video form showing the model supported option

Resolution is filtered to what the selected model can produce.

7

Set the output resolution

Only the resolutions the selected model supports appear here, so a model with a single fixed output shows one option and a model spanning a range shows the full set.

Matching output resolution to the source frame is worth a moment of thought. Animating a small image at a high output resolution asks the model to invent detail that was never in the input, and the result usually reads softer than the still it came from.

Models available for image to video

Thirteen models. This list adds the Vidu Q3 Drama and Turbo variants that text to video does not carry, and several models expose different limits here than they do there.

ModelBest forNotes
Gemini Omni Flash
Default
First passesThe default. 720p, 3 to 10 seconds. Fast enough to test whether a frame animates well.
Vidu Q3 DramaLong single takes720p to 1080p, 2 to 30 seconds. The longest generation available in the workspace.
Vidu Q3 TurboSpeed on the Vidu family540p to 1080p, 1 to 16 seconds. The fast variant beside Q3 Pro.
Vidu Q3 ProQuality on the Vidu family540p to 1080p, 1 to 16 seconds. Accepts clips as short as one second.
Seedance 2.0 MiniReaching 4K480p to 4K, 4 to 15 seconds. Its ceiling here is higher than in text to video.
MiniMax H3Consistent 2K output2K, 5 to 15 seconds.
Luma Ray 3.2Motion quality540p to 1080p, at fixed 5 or 10 second lengths.
Happy Horse 1.1General animation720p to 1080p, 3 to 15 seconds.
Happy Horse 1.0Comparison against 1.1Same range as the newer version, useful when 1.1 misreads a frame.
Seedance 2.0 VIP 1080pGuaranteed 1080p1080p fixed, 4 to 15 seconds.
Seedance 2.0 UncensoredSubjects other models refuse720p, at 5, 10 or 15 seconds.
Google Veo 3.1Prompt adherenceStrong on following a motion instruction literally.
Kling 3.0 PROPrompt control and sound3 to 15 seconds, with a negative prompt field and audio generation.

Resolution and duration figures come from the model picker in the workspace, which prints them beside each model name. Read them in the form you are actually using, because limits are set per generation type rather than per model: Seedance 2.0 Mini reaches 4K in image to video and stops at 1080p in text to video.

Settings reference

Every control in the image to video form, what it changes and the values it accepts.

Image Source

The first frame. Accepts an upload, one of your previous completed tasks, a saved AI Actor, a saved AI Character, or a direct image URL.

JPEG, PNG, WEBP, GIF, AVIF up to 100 MB, minimum 300 by 300 pixels

Last Frame Image

Optional final frame, taking the same five input methods. Supplying it pins where the clip ends instead of leaving that to the model.

Optional, same formats as the first frame

Detailed Prompt

Optional here, since the model already has the scene. Best spent entirely on what moves and what the camera does.

Up to 2,500 or 4,000 characters, depending on the model

Model

Decides the resolution ceiling, the duration range and which extra controls exist. The image to video list is longer than the text to video one.

Thirteen models, including the Vidu Q3 Drama and Turbo variants

Duration

Clip length in seconds, bounded by the selected model. A slider where the model accepts a range, discrete values where it does not.

1 to 30 seconds across the roster

Video resolution

Output resolution, filtered to the selected model. Raising it well above the resolution of the source image asks the model to invent detail.

480p through 4K, model dependent

Aspect ratio

The frame shape of the output. Where it differs from the source image the model has to decide what happens at the edges, so matching the source is the safer default.

Model dependent

Choosing a frame that animates well

Some stills animate cleanly and some fight it, and the difference is usually visible before you generate anything. A frame with clear depth, meaning a foreground subject separated from a background, gives the model somewhere to put parallax when the camera moves. A flat frame with everything on one plane gives it nothing, and the result often slides rather than moves.

Hands, faces at small scale in frame and text on signage are the usual trouble spots. They are stable in a still because nothing is asking them to change, and they tend to deform once motion is applied. Where the shot allows it, framing so those elements are larger and better lit costs nothing and removes most of the problem.

The aspect ratio of the source is worth matching in the output. Where the two differ the model has to decide what exists beyond the edges of your image, and inventing that border is a different task from animating what you gave it.

When to fill the last frame field

Leaving the last frame empty is the right default. The model resolves the ending itself, which is what you want when the shot is exploratory and you are still finding out what the movement should be.

Filling it earns its place when the clip has to hand off to something else. A shot that cuts to a second clip reads better when its final frame is close to that clip's opening frame, and pinning the ending is the direct way to get there. The same applies to a loop, where the last frame needs to return to the first.

Where both end points genuinely matter more than the path between them, the Start-End Frame Video type is the better tool. It is built around that constraint rather than treating it as an option, and it exposes a movement amplitude control that image to video does not.

Model tutorials

Deeper coverage of the models in this roster, with sample output and the frames where each one struggles.

Image to video questions

Do I need a prompt for image to video?

No, the prompt is optional. The model already has the scene from your image. When you do write one, spend it entirely on movement, meaning what travels, in which direction, how fast, and what the camera does. Describing what is already visible in the frame adds nothing.

The drop zone takes JPEG, PNG, WEBP, GIF and AVIF up to 100 MB, with a minimum of 300 by 300 pixels. You can also paste a direct image URL, pick one of your previous completed tasks, or select a saved AI Actor or AI Character instead of uploading anything.

Image to Video animates from a first frame, and the last frame field is an optional constraint on that animation. Start-End Frame Video is built the other way round: two images are the required input and generating the transition between them is the whole job. Use image to video when you care about the movement, and start and end frame when you care about both end points.

Yes, and without downloading it first. Choose Previous Task in the Image Source block and pick the completed generation. That is the normal way to chain a still into a clip, and it keeps both results in the same Artifact Gallery.

Vidu Q3 Drama, at up to 30 seconds. It is the only model in the workspace that reaches that length in a single generation, and it is available in image to video rather than text to video. Most other models stop between 10 and 16 seconds.

Limits are set per generation type, not per model name. Seedance 2.0 Mini reads 480p to 4K in the image to video list and 480p to 1080p in the text to video one. The picker prints the current range beside the model name, so read it in the form you are actually using.

Roughly, yes. Asking for an output well above the resolution of the input means the model is inventing detail that was never there, and the clip often reads softer than the still it came from. Where you need both length and resolution, upscale the finished clip afterwards instead.

Most video models charge per second of output, and higher resolutions cost more, so the figure moves with the model, duration and resolution. The exact cost is shown in the form before you submit, and failed tasks are refunded automatically during credit reconciliation.

Other ways to make a video

Image to Video is one of six routes into the video workspace.