EditorNodesPricingBlog

TEXT TO VIDEO

Write the shot, then run it

Describe a scene and get a clip back from any of thirteen models. Set the duration, pick from six aspect ratios, and read each model's resolution and length limits before you commit to it. No GPU of your own and nothing to install.

Updated:

What text to video does

Text to Video builds a moving clip from a written description with no input file. It is the starting point for a shot that does not exist yet, and the only generation type in the video workspace where the model decides the first frame as well as the movement.

That freedom is also the difficulty. Composition, wardrobe, lighting, camera position and motion are all being resolved in a single pass, and a prompt that names only the subject leaves five of those six to chance. The skill is deciding which of them you actually care about and writing those down.

The text to video form, control by control

Six controls decide almost every outcome here, and the model picker quietly constrains three of the others.

Generation type dropdown in the video workspace with Text to Video selected

The generation type dropdown. Changing the type rebuilds the form beneath it.

1

Pick the generation type

The video workspace opens on a type dropdown holding five entries. Image to Video, Text to Video, Start-End Frame Video, Draw to Video and Video Enhancer. Text to Video is the only one that needs no input file at all.

Changing the type here swaps the form rather than navigating away, so moving between generating a shot from scratch and animating a frame you already have does not cost you your place.

2

Choose a model, and read the numbers beside it

This dropdown carries more information than most model pickers. Each row shows the resolution ceiling and the duration range that model accepts, so MiniMax H3 reads 2K and 5s to 15s, while Gemini Omni Flash reads 720p and 3s to 10s. You can rule a model out before selecting it.

Those numbers are constraints rather than suggestions. Picking a model with a 720p ceiling means the resolution control below will offer you 720p and nothing higher, so if the delivery needs 1080p, the decision was already made here.

Video model dropdown listing each model with its resolution ceiling and duration range

The model picker prints the resolution range and duration range beside every name, before you select it.

Detailed Prompt field in the text to video form with a live character counter

The prompt field, with a character counter that reflects the selected model.

3

Describe the motion, not just the scene

The counter under the field shows the ceiling for the model you picked, which runs to 4,000 characters on some and 2,500 on others. That is far more room than a single sentence, and on video the extra room should go on movement.

A prompt that describes only the scene leaves the model to invent the camera. Naming what moves and how, meaning the direction of a pan, the speed of a push in, whether the subject or the camera is the thing travelling, is the difference between a clip you can cut with and a clip that drifts.

4

Set the duration, and start short

Duration is a slider whose end points come from the model. On the model shown it runs from 5 to 15 seconds, and it starts at the low end for a reason.

Video models bill per second of output, so duration is the setting with the most direct effect on what a mistake costs. Establish that a prompt produces the movement you wanted at the shortest length available, then raise it.

Duration slider in the text to video form running from 5 to 15 seconds

The duration slider. Its range comes from the selected model.

Aspect ratio selector showing ultra wide, landscape, standard, square and portrait options

Six aspect ratios, each drawn as the frame shape it produces.

5

Choose the frame shape

Six ratios are offered, labelled by use rather than by number alone. 21:9 ultra wide, 16:9 landscape, 4:3 standard, 1:1 square, 3:4 portrait and 9:16 portrait.

Aspect ratio is a composition decision taken before generation, not a crop applied after. The model stages the subject differently in a 9:16 vertical frame than in a 21:9 one, so a vertical cut and a widescreen cut of the same idea are two generations rather than two exports.

6

Set the output resolution

The resolution control only offers what the selected model can produce. On a model with a single fixed output it shows one option, and on a model spanning 540p to 1080p it shows the full set.

Resolution and duration together drive both the wait and the cost, and they compound. A 15 second clip at the top of a model range takes considerably longer to return than a 5 second test at the bottom of it.

Video resolution control in the text to video form showing the single option a model allows

Resolution options are filtered to what the selected model supports.

Models available for text to video

Thirteen models, differing on resolution ceiling, duration range, speed and whether they can produce sound. The picker in the workspace shows the first two of those beside each name.

ModelBest forNotes
Gemini Omni Flash
Default
First passes and prompt testingThe default. 720p, 3 to 10 seconds. Fast enough to use as the draft stage.
MiniMax H3Resolution without a long clip2K output, 5 to 15 seconds. The highest resolution ceiling in the text to video list.
Luma Ray 3.2Motion quality540p to 1080p, at fixed 5 or 10 second lengths rather than a free range.
Vidu Q3 ProUnusually short or long clips540p to 1080p, 1 to 16 seconds. The widest duration range offered here.
Seedance 2.0 MiniCheap iteration480p to 1080p, 4 to 15 seconds. The compact member of the Seedance family.
Seedance 2.0 VIP 1080pDelivery at 1080p1080p fixed, 4 to 15 seconds. The Seedance variant that does not drop below 1080p.
Seedance 2.0 VIPSeedance quality at 720p720p, at 5, 10 or 15 seconds. Discrete lengths rather than a slider range.
Seedance 2.0 UncensoredSubjects other models refuse720p, at 5, 10 or 15 seconds.
Happy Horse 1.1General purpose shots720p to 1080p, 3 to 15 seconds. The newer of the two Happy Horse models.
Happy Horse 1.0Comparison against 1.1Same 720p to 1080p range and 3 to 15 second span as the newer version.
Kling 3.0 PROPrompt control and sound3 to 15 seconds, with a negative prompt field and audio generation.
Google Veo 3.1Prompt adherenceGoogle's video model, strong on following an instruction literally.
LTX 2.3Fast turnaroundThe LTX video model, aimed at speed over maximum fidelity.

Resolution and duration figures come from the model picker in the workspace, which prints them beside each model name. The controls in the form change with the model too. A negative prompt field or an audio toggle appears only where the selected model supports it, so confirm the setting you depend on is present before building a workflow around it.

Settings reference

Every control in the text to video form, what it changes and the values it accepts.

Detailed Prompt

The description the model generates from. On video this has to cover movement as well as subject, because anything you leave unstated about the camera gets invented.

Up to 2,500 or 4,000 characters, depending on the model

Model

Decides the resolution ceiling, the duration range, and whether a negative prompt or an audio toggle exists at all. The picker shows both ranges beside each name.

Thirteen models in the text to video list

Duration

Clip length in seconds. A slider on models with a continuous range, and discrete choices on models that only accept specific lengths.

1 to 30 seconds across the roster, commonly 3 to 15

Aspect ratio

The frame shape. Changes how the model composes the shot, so it belongs before generation rather than in an edit afterwards.

21:9, 16:9, 4:3, 1:1, 3:4, 9:16

Video resolution

Output resolution, filtered to what the selected model supports. Together with duration it is the main driver of both wait time and cost.

480p through 2K, model dependent

Negative prompt

A list of elements to keep out of the clip. Appears only on models that accept it, Kling 3.0 PRO among them, and works better than phrasing exclusions inside the main prompt.

Model dependent

Audio

Generates a soundtrack alongside the picture on models that support it. Absent from the form entirely where the model has no audio capability.

Toggle, model dependent

Writing prompts that describe movement

A prompt written for an image describes a moment. A prompt written for video has to describe a change over time, and the most common failure here is a perfectly good still image description handed to a video model, which then invents a camera move to fill the gap.

Compare "a woman standing at a crossing at night" against "a woman in a dark wool coat stands still at a crossing, wet asphalt, sodium street lights behind her, the camera pushes in slowly from a wide to a medium over the full clip, traffic passes left to right in the background". The second names what moves, what stays still, the direction of the camera travel and its speed relative to the clip length.

Stating what holds still is as useful as stating what moves. Video models tend toward constant motion, so a shot that needs a locked camera generally has to say so. The same applies to the subject. If the person should not walk, the prompt has a better chance if it says they are standing.

Where the model offers a negative prompt, that field is where exclusions belong. Telling the main prompt to avoid something tends to produce it, whereas listing it as an exclusion works as intended on the models that accept the field.

Reading the model picker before you generate

The two figures printed beside every model name are the most useful thing in this form. They tell you the resolution ceiling and the duration range, and both of them constrain controls further down. Choosing a 720p model and then looking for a 1080p option is a wasted step, because that option was removed the moment the model was selected.

The ranges also differ in kind. Some models expose a continuous slider, so any length between the end points is available. Others offer only discrete lengths, which is why Seedance 2.0 VIP reads 5s, 10s and 15s rather than a span. A workflow that assumes 7 seconds is available will fail on the second kind.

Working down the list from the cheapest usable model is the reliable pattern. Gemini Omni Flash returns quickly and caps at 720p and 10 seconds, which makes it the right place to find out whether a prompt works at all. Moving to MiniMax H3 for 2K, or to Seedance 2.0 VIP 1080p for a guaranteed 1080p delivery, is a decision worth making after the movement is settled rather than before.

Model tutorials

Deeper coverage of individual models in this roster, with sample output and the cases where each one struggles.

Text to video questions

Which text to video model should I start with?

Gemini Omni Flash is the default and the sensible first pass, because it returns quickly and is capped at 720p and 10 seconds, which keeps the cost of a wrong prompt low. Once the movement reads correctly, regenerate on a model with the resolution and duration the delivery actually needs.

The range belongs to the model rather than the workspace. The model picker prints each duration range beside the model name, so if you need 16 seconds the choice is made there rather than in the slider. Vidu Q3 Pro reaches 16 seconds in text to video, and most other models stop at 15.

The counter under the prompt field shows the ceiling for the selected model, which is 4,000 characters on some and 2,500 on others. Length only helps when the extra words carry information about motion, camera behaviour, lighting or lens character. Repeating quality adjectives changes nothing.

The form is built from what the selected model supports. A negative prompt field and an audio toggle appear only where the model accepts them, and the resolution control lists only the resolutions that model can produce. If a workflow depends on a specific setting, confirm it is present on the model you plan to use before building around it.

On models that support it, yes. Kling 3.0 PRO generates audio alongside the picture. Where a model has no audio capability the toggle is not shown, so the presence of the control is itself the answer for the model you have selected.

No. Aspect ratio changes how the model stages the shot rather than just the canvas shape, so a 9:16 generation places its subject differently from a cropped widescreen frame. Choose the ratio the delivery needs before generating.

Most video models charge per second of output, and higher resolutions cost more, so the figure moves with the model, the duration and the resolution. The exact cost is shown in the form before you submit. Failed tasks are refunded automatically during credit reconciliation.

Typically 30 seconds to 5 minutes. It scales with model, duration and resolution, so a short low resolution test returns far faster than a long clip at the top of a model range.

Other ways to make a video

Text to Video is one of six routes into the video workspace.