TEXT TO VIDEO
Write the shot, then run it
Describe a scene and get a clip back from any of thirteen models. Set the duration, pick from six aspect ratios, and read each model's resolution and length limits before you commit to it. No GPU of your own and nothing to install.
Updated:
What text to video does
Text to Video builds a moving clip from a written description with no input file. It is the starting point for a shot that does not exist yet, and the only generation type in the video workspace where the model decides the first frame as well as the movement.
That freedom is also the difficulty. Composition, wardrobe, lighting, camera position and motion are all being resolved in a single pass, and a prompt that names only the subject leaves five of those six to chance. The skill is deciding which of them you actually care about and writing those down.
The text to video form, control by control
Six controls decide almost every outcome here, and the model picker quietly constrains three of the others.

The generation type dropdown. Changing the type rebuilds the form beneath it.
Pick the generation type
The video workspace opens on a type dropdown holding five entries. Image to Video, Text to Video, Start-End Frame Video, Draw to Video and Video Enhancer. Text to Video is the only one that needs no input file at all.
Changing the type here swaps the form rather than navigating away, so moving between generating a shot from scratch and animating a frame you already have does not cost you your place.
Choose a model, and read the numbers beside it
This dropdown carries more information than most model pickers. Each row shows the resolution ceiling and the duration range that model accepts, so MiniMax H3 reads 2K and 5s to 15s, while Gemini Omni Flash reads 720p and 3s to 10s. You can rule a model out before selecting it.
Those numbers are constraints rather than suggestions. Picking a model with a 720p ceiling means the resolution control below will offer you 720p and nothing higher, so if the delivery needs 1080p, the decision was already made here.

The model picker prints the resolution range and duration range beside every name, before you select it.

The prompt field, with a character counter that reflects the selected model.
Describe the motion, not just the scene
The counter under the field shows the ceiling for the model you picked, which runs to 4,000 characters on some and 2,500 on others. That is far more room than a single sentence, and on video the extra room should go on movement.
A prompt that describes only the scene leaves the model to invent the camera. Naming what moves and how, meaning the direction of a pan, the speed of a push in, whether the subject or the camera is the thing travelling, is the difference between a clip you can cut with and a clip that drifts.
Set the duration, and start short
Duration is a slider whose end points come from the model. On the model shown it runs from 5 to 15 seconds, and it starts at the low end for a reason.
Video models bill per second of output, so duration is the setting with the most direct effect on what a mistake costs. Establish that a prompt produces the movement you wanted at the shortest length available, then raise it.

The duration slider. Its range comes from the selected model.

Six aspect ratios, each drawn as the frame shape it produces.
Choose the frame shape
Six ratios are offered, labelled by use rather than by number alone. 21:9 ultra wide, 16:9 landscape, 4:3 standard, 1:1 square, 3:4 portrait and 9:16 portrait.
Aspect ratio is a composition decision taken before generation, not a crop applied after. The model stages the subject differently in a 9:16 vertical frame than in a 21:9 one, so a vertical cut and a widescreen cut of the same idea are two generations rather than two exports.
Set the output resolution
The resolution control only offers what the selected model can produce. On a model with a single fixed output it shows one option, and on a model spanning 540p to 1080p it shows the full set.
Resolution and duration together drive both the wait and the cost, and they compound. A 15 second clip at the top of a model range takes considerably longer to return than a 5 second test at the bottom of it.

Resolution options are filtered to what the selected model supports.
Models available for text to video
Thirteen models, differing on resolution ceiling, duration range, speed and whether they can produce sound. The picker in the workspace shows the first two of those beside each name.
| Model | Best for | Notes |
|---|---|---|
| Gemini Omni Flash Default | First passes and prompt testing | The default. 720p, 3 to 10 seconds. Fast enough to use as the draft stage. |
| MiniMax H3 | Resolution without a long clip | 2K output, 5 to 15 seconds. The highest resolution ceiling in the text to video list. |
| Luma Ray 3.2 | Motion quality | 540p to 1080p, at fixed 5 or 10 second lengths rather than a free range. |
| Vidu Q3 Pro | Unusually short or long clips | 540p to 1080p, 1 to 16 seconds. The widest duration range offered here. |
| Seedance 2.0 Mini | Cheap iteration | 480p to 1080p, 4 to 15 seconds. The compact member of the Seedance family. |
| Seedance 2.0 VIP 1080p | Delivery at 1080p | 1080p fixed, 4 to 15 seconds. The Seedance variant that does not drop below 1080p. |
| Seedance 2.0 VIP | Seedance quality at 720p | 720p, at 5, 10 or 15 seconds. Discrete lengths rather than a slider range. |
| Seedance 2.0 Uncensored | Subjects other models refuse | 720p, at 5, 10 or 15 seconds. |
| Happy Horse 1.1 | General purpose shots | 720p to 1080p, 3 to 15 seconds. The newer of the two Happy Horse models. |
| Happy Horse 1.0 | Comparison against 1.1 | Same 720p to 1080p range and 3 to 15 second span as the newer version. |
| Kling 3.0 PRO | Prompt control and sound | 3 to 15 seconds, with a negative prompt field and audio generation. |
| Google Veo 3.1 | Prompt adherence | Google's video model, strong on following an instruction literally. |
| LTX 2.3 | Fast turnaround | The LTX video model, aimed at speed over maximum fidelity. |
Resolution and duration figures come from the model picker in the workspace, which prints them beside each model name. The controls in the form change with the model too. A negative prompt field or an audio toggle appears only where the selected model supports it, so confirm the setting you depend on is present before building a workflow around it.
Settings reference
Every control in the text to video form, what it changes and the values it accepts.
Detailed Prompt
The description the model generates from. On video this has to cover movement as well as subject, because anything you leave unstated about the camera gets invented.
Up to 2,500 or 4,000 characters, depending on the model
Model
Decides the resolution ceiling, the duration range, and whether a negative prompt or an audio toggle exists at all. The picker shows both ranges beside each name.
Thirteen models in the text to video list
Duration
Clip length in seconds. A slider on models with a continuous range, and discrete choices on models that only accept specific lengths.
1 to 30 seconds across the roster, commonly 3 to 15
Aspect ratio
The frame shape. Changes how the model composes the shot, so it belongs before generation rather than in an edit afterwards.
21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Video resolution
Output resolution, filtered to what the selected model supports. Together with duration it is the main driver of both wait time and cost.
480p through 2K, model dependent
Negative prompt
A list of elements to keep out of the clip. Appears only on models that accept it, Kling 3.0 PRO among them, and works better than phrasing exclusions inside the main prompt.
Model dependent
Audio
Generates a soundtrack alongside the picture on models that support it. Absent from the form entirely where the model has no audio capability.
Toggle, model dependent
Writing prompts that describe movement
A prompt written for an image describes a moment. A prompt written for video has to describe a change over time, and the most common failure here is a perfectly good still image description handed to a video model, which then invents a camera move to fill the gap.
Compare "a woman standing at a crossing at night" against "a woman in a dark wool coat stands still at a crossing, wet asphalt, sodium street lights behind her, the camera pushes in slowly from a wide to a medium over the full clip, traffic passes left to right in the background". The second names what moves, what stays still, the direction of the camera travel and its speed relative to the clip length.
Stating what holds still is as useful as stating what moves. Video models tend toward constant motion, so a shot that needs a locked camera generally has to say so. The same applies to the subject. If the person should not walk, the prompt has a better chance if it says they are standing.
Where the model offers a negative prompt, that field is where exclusions belong. Telling the main prompt to avoid something tends to produce it, whereas listing it as an exclusion works as intended on the models that accept the field.
Reading the model picker before you generate
The two figures printed beside every model name are the most useful thing in this form. They tell you the resolution ceiling and the duration range, and both of them constrain controls further down. Choosing a 720p model and then looking for a 1080p option is a wasted step, because that option was removed the moment the model was selected.
The ranges also differ in kind. Some models expose a continuous slider, so any length between the end points is available. Others offer only discrete lengths, which is why Seedance 2.0 VIP reads 5s, 10s and 15s rather than a span. A workflow that assumes 7 seconds is available will fail on the second kind.
Working down the list from the cheapest usable model is the reliable pattern. Gemini Omni Flash returns quickly and caps at 720p and 10 seconds, which makes it the right place to find out whether a prompt works at all. Moving to MiniMax H3 for 2K, or to Seedance 2.0 VIP 1080p for a guaranteed 1080p delivery, is a decision worth making after the movement is settled rather than before.
Model tutorials
Deeper coverage of individual models in this roster, with sample output and the cases where each one struggles.
Text to video questions
Why does the duration slider stop before the length I want?
The range belongs to the model rather than the workspace. The model picker prints each duration range beside the model name, so if you need 16 seconds the choice is made there rather than in the slider. Vidu Q3 Pro reaches 16 seconds in text to video, and most other models stop at 15.
How long can a text to video prompt be?
The counter under the prompt field shows the ceiling for the selected model, which is 4,000 characters on some and 2,500 on others. Length only helps when the extra words carry information about motion, camera behaviour, lighting or lens character. Repeating quality adjectives changes nothing.
Why does the form change when I switch model?
The form is built from what the selected model supports. A negative prompt field and an audio toggle appear only where the model accepts them, and the resolution control lists only the resolutions that model can produce. If a workflow depends on a specific setting, confirm it is present on the model you plan to use before building around it.
Can I generate sound with the video?
On models that support it, yes. Kling 3.0 PRO generates audio alongside the picture. Where a model has no audio capability the toggle is not shown, so the presence of the control is itself the answer for the model you have selected.
Should I generate at 16:9 and crop to vertical?
No. Aspect ratio changes how the model stages the shot rather than just the canvas shape, so a 9:16 generation places its subject differently from a cropped widescreen frame. Choose the ratio the delivery needs before generating.
What does a text to video generation cost?
Most video models charge per second of output, and higher resolutions cost more, so the figure moves with the model, the duration and the resolution. The exact cost is shown in the form before you submit. Failed tasks are refunded automatically during credit reconciliation.
How long does the generation take?
Typically 30 seconds to 5 minutes. It scales with model, duration and resolution, so a short low resolution test returns far faster than a long clip at the top of a model range.
Other ways to make a video
Text to Video is one of six routes into the video workspace.
Video & LipSync
- Video Generator
- Text to Video
- Image to Video
- Start-End Frame to Video
- Draw to Video
- Motion Control
- Video Enhancer
- Video Upscaler
- Video to Video LipSync
- Audio to Video LipSync
- Image to Video LipSync
- Video FaceSwap
- Seedance 2
- Vidu Q3 Pro
- Gemini Omni
- Google Veo 3.1
- Kling 3.0 Pro
- Luma Ray 3.2
- LTX 2.3
- Happy Horse 1.1
- WAN 2.7
- Kling 3.0 Motion
- ByteDance Upscaler
- InfiniteTalk
- InsightFace