DRAW TO VIDEO
Point at the motion instead of describing it
Load an image, draw arrows where things should move, number them to set the order, and generate. Arrow length sets speed, and the annotations are cleared from the clip after the first frame.
Updated:
What Draw to Video does
Draw to Video gives you a canvas over an image and reads what you draw on it as motion instructions. Your drawing is converted to an image and processed as image to video, so the underlying job is the same as animating a still. What changes is how you tell the model what should move.
Spatial instructions are the case this solves. Describing three elements moving in three different directions takes a paragraph of careful prose that the model may still misread. Drawing three arrows takes ten seconds and leaves no ambiguity about which element each instruction applies to.
The Draw to Video canvas
One surface holding the drawing tools, the base image, the model and the prompt together.

Draw to Video sits fourth in the generation type list.
Open Draw to Video
Draw to Video is the fourth entry in the generation type dropdown, between Start-End Frame Video and Video Enhancer. Selecting it opens a canvas rather than rebuilding the standard form.
That canvas is the difference. Every other generation type takes text and files, and this one takes a drawing you make over an image.
Load an image and read the tips panel
The canvas needs a base image before you can draw on it. The drop zone accepts JPEG, PNG, WEBP, GIF and AVIF up to 300 MB at a minimum of 300 by 300 pixels, and the interface will not let you start drawing without one.
The Drawing Tips panel on the right states the grammar the model reads. Arrows indicate motion direction, text with numbers sequences that motion, and longer arrows mean faster or larger movement. Those three rules are the whole interface between your drawing and the result.

The canvas takes the left side, the model, prompt and drawing tips sit on the right. Cost appears in the form before you submit.

The toolbar along the bottom of the canvas, with undo, redo and clear at the right.
Annotate with the drawing tools
The toolbar holds a select tool, an arrow tool, a freehand brush, shape tools, a text tool, an image tool and a fill tool, with undo, redo and two clear actions grouped at the right hand end.
The arrow and text tools carry most of the weight, because they are the two the model reads as instructions. The rest are there for marking areas and adding context around them.
The model behind the canvas
Draw to Video runs on a single model rather than offering the full video roster.
| Model | Best for | Notes |
|---|---|---|
| Happy Horse 1.0 Default | Reading sketched motion instructions | The model behind the canvas. 720p to 1080p, 3 to 15 seconds, with its own options behind a Configure button beside the picker. |
Happy Horse 1.0 also appears in the standard image to video list, where Happy Horse 1.1 sits beside it. If you want the newer version, animate the image through Image to Video and describe the movement in text instead of drawing it.
Settings reference
Everything the canvas exposes, what it changes and the values it accepts.
Base image
The picture you draw over. Required before the drawing tools become active, and it supplies the first frame of the generated clip.
JPEG, PNG, WEBP, GIF, AVIF up to 300 MB, minimum 300 by 300 pixels
Arrows
Read as motion direction. Length carries meaning as well as angle, since a longer arrow indicates faster or larger movement.
Drawn on the canvas with the arrow tool
Numbered text
Read as sequence. Adding text with numbers tells the model the order in which the marked movements should happen.
Drawn on the canvas with the text tool
Prompt
Arrives with an instruction already in it, telling the model to remove the arrows, text and annotations after the first frame. Keep that instruction and add your own description around it.
Up to 2,500 characters
Model
Happy Horse 1.0, with additional settings behind a Configure button beside the picker rather than laid out in the form.
720p to 1080p, 3 to 15 seconds
Canvas tools
Select, arrow, brush, shapes, text, image and fill, with undo, redo and clear. The arrow and text tools are the two the model interprets as instructions.
Toolbar along the bottom of the canvas
The drawing grammar
Three conventions carry every instruction here, and the tips panel states all of them. An arrow indicates the direction something moves. Text containing a number sets the order in which the marked movements happen. And the length of an arrow indicates how fast or how large that movement should be.
Arrow length being meaningful is the part people miss. A short arrow and a long arrow pointing the same way are two different instructions, so drawing a long sweeping arrow for a subtle drift will produce more movement than you wanted. Sizing the arrow to the movement is the habit worth building.
Marks outside these three conventions do not become instructions. Shading an area or circling a subject with the brush marks the picture without telling the model to do anything with it, so the arrow and text tools are where the useful work happens.
Why the prompt starts with an instruction in it
The prompt field opens carrying a sentence asking the model to remove arrows, text, annotations, drawings, labels, symbols and markings after the first frame. It is there because the annotated image is what gets sent, so without that instruction your arrows would be treated as part of the picture and stay visible through the clip.
The interface says explicitly to keep it, and a note under the field repeats that. Add your own description around the existing sentence rather than clearing the field and starting fresh. The counter runs to 2,500 characters, and that opening instruction uses fewer than a hundred of them.
Anything you add should cover what the drawing cannot. The arrows handle direction, speed and sequence, which leaves lighting, atmosphere, the behaviour of elements you did not annotate, and the overall feel of the shot as the useful things to write about.
When pointing beats describing
Text is good at describing one thing well and poor at describing where several things are. A prompt reading "the car moves left while the crowd moves right and the camera pushes in" asks the model to work out which crowd, which car and how far, from a sentence that contains none of that spatial information.
Three arrows answer all of it at once. Each one is attached to the thing it applies to, its angle carries the direction and its length carries the amount. Nothing has to be matched up by name, because the instruction is sitting on top of the object it governs.
The reverse also holds. A shot where one subject moves in one obvious way is faster to describe in a sentence than to annotate, and Image to Video is the better route for it. Reach for the canvas when the instruction is spatial, and for the prompt field when it is not.
Getting a base image worth annotating
The canvas needs a picture before it does anything, and the quality of that picture sets the ceiling on the result. Generating it in the image workspace first is the usual route, and it means the frame can be revised cheaply before any video credit is spent on it.
Frames with visible separation between elements annotate better than flat ones. If the car, the crowd and the background all sit on distinct planes, an arrow drawn over one of them is unambiguous. If everything overlaps, the arrow has to be read against a cluttered region and the instruction gets weaker.
Leave room for the marks. An image composed edge to edge with detail gives you nowhere to draw an arrow that is not sitting on something important, and arrows drawn across a subject are harder for the model to attribute than arrows drawn beside it.
Related tutorials
Coverage of the model behind the canvas and of the alternative routes to animating a still.
Draw to Video questions
Why is the prompt already filled in?
It arrives carrying an instruction to remove arrows, text, annotations, drawings, labels, symbols and markings after the first frame. Without it your annotations would persist into the generated clip as visible marks. The interface says to keep that instruction, so add your own description around it rather than replacing it.
Do I need an image to start, or can I draw from scratch?
You need an image. The canvas prompts you to upload one before drawing becomes available, and it accepts JPEG, PNG, WEBP, GIF and AVIF up to 300 MB at a minimum of 300 by 300 pixels. The drawing is an annotation layer over that picture rather than a replacement for it.
How is this different from image to video?
Both animate a still. Image to Video takes the motion instruction as text, and Draw to Video takes it as marks placed directly on the frame. Where the movement is easier to point at than to describe, such as three elements moving in different directions, the drawing carries the spatial information a sentence would struggle with.
Which model does Draw to Video use?
Happy Horse 1.0, at 720p to 1080p and 3 to 15 seconds. Additional model settings are available behind a Configure button beside the picker rather than being laid out in the form.
How do I sequence several movements?
Add text containing numbers next to each arrow. The Drawing Tips panel states that numbered text sequences motion, so labelling one arrow 1 and another 2 tells the model which movement should happen first.
How do I control the speed of a movement?
Arrow length. A longer arrow indicates faster or larger movement, so the same direction drawn short and drawn long produces different results. This is the one place where the size of a mark carries meaning rather than just its position.
What does a Draw to Video generation cost?
The credit cost appears in the form before you click Create, and it depends on the model and the settings you chose. Failed tasks are refunded automatically during credit reconciliation.
Other ways to make a video
Draw to Video is one of six routes into the video workspace.
Video & LipSync
- Video Generator
- Text to Video
- Image to Video
- Start-End Frame to Video
- Draw to Video
- Motion Control
- Video Enhancer
- Video Upscaler
- Video to Video LipSync
- Audio to Video LipSync
- Image to Video LipSync
- Video FaceSwap
- Seedance 2
- Vidu Q3 Pro
- Gemini Omni
- Google Veo 3.1
- Kling 3.0 Pro
- Luma Ray 3.2
- LTX 2.3
- Happy Horse 1.1
- WAN 2.7
- Kling 3.0 Motion
- ByteDance Upscaler
- InfiniteTalk
- InsightFace