TEXT TO IMAGE
Describe it, then generate it
Write a prompt and get an image back from any of eight models. Control the aspect ratio, hold a seed to iterate deliberately, and apply your own trained style. Nothing to install and no GPU of your own.
What text to image does
Text to Image builds a picture from a written description with no input file. It is the starting point for concept work, storyboard frames, character exploration and any shot that does not exist yet to photograph or edit.
The useful skill is not writing longer prompts. It is knowing which model to reach for, and using the seed to turn each generation into a test of one specific change rather than another roll of the dice.
The text to image form, control by control
Six controls decide almost every outcome here. This is what each one does and when it matters.

The generation type dropdown. Selecting a type rebuilds the form beneath it.
Pick the generation type
The image workspace opens on a type dropdown listing Text to Image, Image to Image, Image Enhance, Character Generation and Model Training. Text to Image is the one that needs no input file.
Changing the type here swaps the entire form rather than navigating away, so you can move between generating from scratch and editing a result without losing your place.
Choose a model
The model you pick decides more than image style. It also decides which controls the rest of the form shows you, because a negative prompt field, a seed input or a custom dimension option only appears where that model supports it.
Nano Banana 2 Lite is the default and the fastest way to test whether a prompt is working. Move to a heavier model once the composition is settled rather than paying for quality on drafts you will discard.

The model picker. Nano Banana 2 Lite is selected by default.

The prompt field. Most models accept up to 2,500 characters.
Write the prompt
Most models in the roster take prompts up to 2,500 characters, which is far more room than a single sentence. Use it on the things that change the frame, meaning subject, camera position, lens character, lighting direction and time of day.
Where the model supports a negative prompt, that field is the place to exclude what keeps appearing uninvited. Listing unwanted elements there works better than telling the main prompt what you do not want.
Set the aspect ratio
Aspect ratio is a composition decision, not a crop applied afterwards. The model frames the subject differently at 16:9 than at 1:1, so choosing it up front produces a better image than generating square and cropping later.
Pick the ratio your delivery needs. A 9:16 vertical frame for social and a 16:9 frame for a timeline are genuinely different generations rather than two exports of one file.

Aspect ratio options. Some models also accept custom pixel dimensions.

The seed field. Reusing a value reproduces a result you already liked.
Fix the seed when you want to iterate
A seed is the starting noise for a generation. Leave it empty and every run is different. Enter the seed from a result you liked and you get that image back, which turns prompt editing into a controlled comparison.
This is the difference between guessing and testing. Holding the seed steady while changing one clause in the prompt shows you what that clause actually did.
Save a face you want to keep
When a generation produces a face worth reusing, the save actor and save character options store it so later work can call it back. This is the bridge between one good frame and a consistent series.
If you already know a project needs the same person in many shots, Character Generation is the better starting point, since it accepts reference photos directly.

Save actor and save character, for reusing a face across later generations.
Models available for text to image
Eight models, and the choice is a real one. They differ on speed, on how literally they follow a prompt, and on whether output leans photographic or illustrative.
| Model | Best for | Notes |
|---|---|---|
| Nano Banana 2 Lite Default | Drafting and iteration | Fast with consistent quality. The default, and the sensible place to test a prompt. |
| Nano Banana Pro Ultra | Final frames | The highest quality variant in the Nano Banana family. Slower than Lite. |
| GPT Image 2 | Prompt following and text in image | OpenAI image model. Strong on instructions that describe a specific arrangement. |
| FLUX.2 Pro | Detail and photographic realism | Professional FLUX generation, with strong handling of fine texture. |
| HiDream-O1 | Stylised and illustrative work | HiDream image model, a useful alternative when FLUX output looks too photographic. |
| MidJourney 8.0 | Art direction and mood | MidJourney style generation, which favours aesthetic strength over literal accuracy. |
| Krea 2 Large | Quality on complex scenes | The larger of the two Krea models. Use where a scene has many interacting elements. |
| Krea 2 Turbo | Speed | The faster Krea variant, aimed at quick passes rather than final output. |
The controls in the form change with the model. A negative prompt field, a seed input or custom pixel dimensions appear only where the selected model supports them, so check the setting you depend on is present before building a workflow around it.
Settings reference
Every control in the text to image form, what it changes and the values it accepts.
Prompt
The description the model generates from. Specific detail about subject, lighting and composition changes the result far more than adjectives about quality.
Up to 2,500 characters on most models
Negative prompt
A list of elements to exclude. Only appears on models that support it, and works better than phrasing exclusions inside the main prompt.
Model dependent
Aspect ratio
The frame shape. Affects how the model composes the shot, so set it before generating rather than cropping after.
1:1, 16:9, 9:16, 4:3 and others, plus custom dimensions on some models
Number of images
How many variations to produce in one submission. A small batch is the cheapest way to compare directions.
Commonly 1 to 4, on models supporting batch output
Seed
The starting value for the generation. Reuse a seed to reproduce a result, or leave it empty for a different image each run.
Integer, optional
LoRA and strength
Applies a style model on top of the base model. Choose a predefined style, one of your own trained LoRAs, or an external LoRA URL, then set how strongly it applies.
Predefined, trained or external URL
Output format
The file type of the result. PNG preserves detail for further editing, JPEG produces smaller files for review and sharing.
PNG or JPEG
Writing prompts that change the output
A weak prompt describes a category. A strong one describes a specific frame. Compare "a woman in a city at night" against "a woman in a dark wool coat standing at a crossing, wet asphalt, sodium street lights behind her, shot on a 35mm lens at eye level, shallow depth of field". The second names subject, wardrobe, surface, lighting source, lens and camera height. Every one of those is a decision the model would otherwise make for you.
Terms like masterpiece, best quality and highly detailed do almost nothing. They describe an outcome rather than a scene, and the model has no way to act on them. The space they take up is better spent on lighting direction or time of day.
When something unwanted keeps appearing, put it in the negative prompt rather than arguing with the main prompt. Telling a model to avoid text tends to produce text, while listing it as an exclusion works as intended on models that support the field.
Model tutorials
Deeper coverage of the individual models in this roster, with sample output and the cases where each one struggles.
Text to image questions
Why do the form controls change when I switch model?
The form is built from what the selected model supports. A negative prompt field, a seed input or custom pixel dimensions appear only on models that accept them, so switching model can add or remove controls. If a workflow depends on a specific setting, check it is present on the model you plan to use.
How do I reproduce an image I already generated?
Enter the seed value from that result into the seed field and keep the prompt and model the same. Holding the seed constant while changing one part of the prompt is also the reliable way to see what a specific phrase is doing.
Which model should I start with?
Nano Banana 2 Lite is the default because it is fast and consistent, which makes it the cheapest way to find out whether a prompt works at all. Once the composition is right, regenerate on a heavier model such as Nano Banana Pro Ultra or FLUX.2 Pro.
Can I apply my own style to a generation?
Yes. The LoRA picker in the text to image form accepts a predefined style, any LoRA you have trained yourself through Model Training, or an external LoRA URL. Strength controls how strongly it applies over the base model.
Should I generate square and crop to other ratios?
No. Aspect ratio changes how the model composes the shot rather than just the canvas shape, so a 16:9 generation frames its subject differently from a cropped square. Set the ratio you need before generating.
What happens if a generation fails?
Failed tasks are refunded automatically during credit reconciliation, so a failure does not cost you credits.
Other ways to generate images
Text to Image is one of seven generation types in the image workspace.
Video & LipSync
- Video Generator
- Text to Video
- Image to Video
- Start-End Frame to Video
- Draw to Video
- Motion Control
- Video Enhancer
- Video Upscaler
- Video to Video LipSync
- Audio to Video LipSync
- Image to Video LipSync
- Video FaceSwap
- Seedance 2
- Vidu Q3 Pro
- Gemini Omni
- Google Veo 3.1
- Kling 3.0 Pro
- Luma Ray 3.2
- LTX 2.3
- Happy Horse 1.1
- WAN 2.7
- Kling 3.0 Motion
- ByteDance Upscaler
- InfiniteTalk
- InsightFace