MOTION CONTROL
Take the performance, change the performer
Give the model a clip carrying the movement you want and an image of the character who should be doing it. Transfer body motion for up to 30 seconds, or drive a camera move from a still for up to 10.
Updated:
What motion control does
Motion Control re-animates a video using a reference character image and a prompt. The movement comes from footage you supply, and the subject on screen comes from an image you supply separately. The person in the source clip does not appear in the output.
That separation is what makes it different from every other type in the video workspace. Everywhere else, motion is described in words and the model interprets the description. Here the motion already exists as footage, so a specific performance can be reproduced rather than approximated.
The motion control form, control by control
Eight controls, and the orientation setting quietly governs both what transfers and how long the clip can be.

Motion Control is the second task type inside the Video Enhancer.
Open the enhancer and switch to Motion Control
Motion Control sits under Video Enhancer in the generation type dropdown, then behind a second task type dropdown that also holds Video Upscaling. The form defaults to upscaling, so this switch is a step people miss.
The two task types share a form and almost nothing else. Selecting Motion Control replaces the model, adds a required character image input and a prompt field, and changes what the source video is being used for.
Select the model
Kling 3.0 Pro is the only motion control model, labelled in the picker with a 30 second maximum and a high fidelity marker. As with upscaling, there is no comparison to work through.
That 30 second figure is the ceiling in Video Orientation. The Image Orientation mode caps at 10 seconds instead, so the limit printed beside the model name is the best case rather than a fixed number.

One model, labelled with its 30 second ceiling.

The source clip supplies the motion, not the picture.
Supply the reference video
The Video Source block takes an upload, a direct video URL, or one of your previous completed tasks, accepting MP4 up to 1000 MB at a minimum of 300 by 300 pixels.
What this clip contributes is the movement rather than the appearance. Its subject will not appear in the output, so the thing to judge when choosing it is whether the motion is clean, well framed and unambiguous.
Supply the character image
The character image is the subject that will appear in the output. It accepts an upload, a direct image URL or a previous completed task, taking JPEG, PNG, WEBP, GIF and AVIF up to 100 MB at a minimum of 300 by 300 pixels.
This block offers three input methods rather than the five available in image to video. The saved AI Actor and AI Character options are not present here, so a character you want to reuse has to come through as a file, a URL or a previous task.

The character image decides who appears. Three input methods, not five.

The prompt describes the motion you want, up to 2,500 characters.
Describe the motion
The prompt field takes up to 2,500 characters and describes the movement or scene you want produced. It works alongside the reference video rather than replacing it.
Vague prompts produce unpredictable results here more than elsewhere, because the model is reconciling three inputs at once. Being explicit about how the character should move, or what the camera should do, is what keeps that reconciliation on track.
Choose the orientation mode
Character Orientation offers Video Orientation, labelled for complex motions, and Image Orientation, labelled for camera movements. This decides what the model transfers, and it also sets the maximum length.
Video Orientation is for a character performing actions such as walking, gesturing or dancing, and allows up to 30 seconds. Image Orientation applies camera moves like panning, zooming and tracking, and caps at 10 seconds. Front facing character images suit the first mode, and angled shots often work better in the second.

The single most consequential control in this form.

Optional extra elements, added one at a time.
Add optional elements
An optional Elements section lets you add further items to the generation through an Add Element control. It is genuinely optional, and a first pass is usually better run without it.
Adding elements increases the number of things the model is reconciling at once. Establish that the character and the motion combine correctly before introducing anything else into the frame.
Decide about the source audio
A single toggle, Keep Original Sound, preserves the audio from the source video in the output. Where the reference clip carries dialogue or ambience worth keeping, this saves reattaching it later.
Where the source is only there to supply movement and its audio is irrelevant, turning this off avoids carrying sound into an edit that will replace it anyway.

Keep Original Sound carries the audio across from the source clip.
The motion control model
One model, with its capability split across two orientation modes rather than across separate models.
| Model | Best for | Notes |
|---|---|---|
| Kling 3.0 Pro Motion Control Default | Transferring motion onto a different character | The only motion control model. Up to 30 seconds in Video Orientation and 10 seconds in Image Orientation, billed per second of output. |
The Video Enhancer form holds two task types with separate model lists. It opens on Video Upscaling, whose single ByteDance entry does nothing like this, so confirm the task type reads Motion Control before reading the model dropdown.
Settings reference
Every control in the motion control form, what it changes and the values it accepts.
Task Type
Switches the Video Enhancer between its two jobs. The form opens on Video Upscaling, so Motion Control has to be selected explicitly.
Video Upscaling or Motion Control
Video Source
The reference clip supplying the movement. Its subject does not appear in the output, so judge it on the clarity of the motion rather than the picture.
MP4 up to 1000 MB, minimum 300 by 300 pixels
Character Image
The subject that will appear in the output. Three input methods here rather than the five offered in image to video, with no saved AI Actor or AI Character options.
JPEG, PNG, WEBP, GIF, AVIF up to 100 MB
Detailed Prompt
Describes the motion or scene you want. Explicit instructions matter more here than elsewhere, because the model is reconciling a video, an image and text at once.
Up to 2,500 characters
Character Orientation
Decides whether the model transfers body motion or camera motion, and sets the maximum length. The most consequential control in this form.
Video Orientation up to 30 seconds, Image Orientation up to 10 seconds
Elements
Optional additional items added to the generation one at a time. Best left empty until the character and motion combine correctly on their own.
Optional, added individually
Keep Original Sound
Carries the audio from the source clip into the output. Useful when the reference carries dialogue or ambience worth preserving.
Toggle
Choosing the orientation mode
This one dropdown decides more than anything else in the form. Video Orientation reads the source clip as body motion and applies it to your character, which is what you want for a person walking, gesturing or dancing. Image Orientation reads it as camera behaviour and applies that to a largely static subject, which is what you want for a pan, a zoom or a tracking move.
The choice also sets the length ceiling. Video Orientation allows up to 30 seconds and Image Orientation stops at 10, so a long shot forces the first mode regardless of what the footage contains. That is worth checking against the edit before you start.
The character image should match the mode. A front facing photograph gives Video Orientation the information it needs to re-pose a body through a full range of movement. Image Orientation is more forgiving of side and angled shots, because the subject is holding still while the camera does the work.
Preparing the two inputs
The reference video should be trimmed to the movement you actually want. Everything else in it is motion the model will try to reproduce, so a clip containing the gesture plus ten seconds of standing around produces a result containing both. Keeping the source under 30 seconds is the stated guidance, and well under it is usually better.
Judge the reference on the clarity of its motion rather than on how it looks. Framing that keeps the moving subject fully in shot, even lighting that does not lose limbs into shadow, and a single unambiguous action all matter more than resolution or grade, since none of the source picture survives into the output.
The character image carries the opposite burden. Everything visible in the result comes from it, so a clean, well lit, reasonably high resolution image is worth preparing properly. Generating that image through Character Generation in the image workspace first is the reliable way to get one, and it also means the same character can drive several motion control runs.
Related tutorials
Deeper coverage of the model behind this form and of the character work that feeds it.
Motion control questions
What is the difference between the two orientation modes?
Video Orientation transfers complex motions such as walking, gesturing or dancing, and allows up to 30 seconds. Image Orientation applies camera movements such as panning, zooming and tracking, and caps at 10 seconds. Front facing character images tend to suit Video Orientation, and side or angled shots often produce better results in Image Orientation.
Why is the maximum length different from the 30 seconds on the model?
The 30 second figure printed beside Kling 3.0 Pro is the ceiling in Video Orientation. Selecting Image Orientation lowers the maximum to 10 seconds, so the number on the model is the best case rather than a fixed limit.
How long should the reference video be?
Under 30 seconds gives the best results, and shorter is usually better still. The clip only needs to contain the movement you want transferred, so trimming it to that section before uploading removes motion the model would otherwise try to reproduce.
What makes a good character image?
A clear, well lit image of the subject. Front facing images work best for Video Orientation, where the body has to be re-posed through a full motion. Side or angled shots can work better for Image Orientation, where the camera is doing the moving rather than the character.
Can I use a saved AI Character here?
Not directly. The character image block offers upload, image URL and previous task, and it does not carry the AI Actor and AI Character options that appear in the image to video form. A saved character has to come in as a previous task, a file or a URL.
How is motion control billed?
Per second of output video. Because the final length is not known at submission, the maximum for your chosen orientation is charged upfront and the difference against the real output length is refunded automatically once the video is generated.
How is this different from image to video?
Image to Video animates a still from a text description of the movement. Motion Control takes the movement from real footage instead, which means you can reproduce a specific performance rather than describing one and hoping. It is the right tool when the motion already exists somewhere.
Other ways to make a video
Motion Control is one of six routes into the video workspace.
Video & LipSync
- Video Generator
- Text to Video
- Image to Video
- Start-End Frame to Video
- Draw to Video
- Motion Control
- Video Enhancer
- Video Upscaler
- Video to Video LipSync
- Audio to Video LipSync
- Image to Video LipSync
- Video FaceSwap
- Seedance 2
- Vidu Q3 Pro
- Gemini Omni
- Google Veo 3.1
- Kling 3.0 Pro
- Luma Ray 3.2
- LTX 2.3
- Happy Horse 1.1
- WAN 2.7
- Kling 3.0 Motion
- ByteDance Upscaler
- InfiniteTalk
- InsightFace