Share this post:
Seedance 2.5 Tutorial: Text to Video and Image to Video
Seedance 2.5 is now available in the AI FILMS Studio video workspace. It is the first model in the roster that reaches 4K and 30 seconds in a single generation. This guide covers both modes it appears in, text-to-video and image-to-video, and the same two workflows in the Nodes Graph Editor.
What Seedance 2.5 Is
Seedance 2.5 is ByteDance's video generation model. Volcano Engine president Tan Dai announced it at the 2026 Volcano Engine FORCE conference in Beijing on 23 June 2026, as the successor to the Seedance 2.0 family.
The headline number is length. Seedance 2.5 produces up to 30 seconds in one pass rather than stitching shorter segments together, which is double the 15 second ceiling of Seedance 2.0. Camera and shot continuity hold across the full clip, so a 30 second generation reads as one take instead of five joined ones.
The second change is how much you can feed it. Seedance 2.5 conditions on up to 50 reference assets across images, video and audio, up from 12 in the previous generation. In the Studio form those 50 slots appear as three separate sections, and the walkthrough below covers each one.
What changed from Seedance 2.0
| Capability | Seedance 2.0 family | Seedance 2.5 |
|---|---|---|
| Maximum duration | 15 seconds | 30 seconds |
| Resolution range | 480p to 1080p | 480p to 4K |
| Reference assets | 12 | 50 |
| Aspect ratios in text-to-video | 4 | 7, including an adaptive option |
| Native audio | Yes | Yes, in the same generation pass |
The model picker prints the range beside the name. Seedance 2.5 reads 480p-4K · 4s-30s, which is the widest span in the video roster.
Text to Video
Text to Video builds a clip from a written description with no input file. Seedance 2.5 is the only entry in this list that can hold a shot for 30 seconds, so it changes what a single generation is useful for.
The text to video workspace. The form sits on the left, the finished clip plays on the right.
The workspace keeps the form and the result side by side. You can adjust a setting and regenerate without losing the previous output.
Everything below happens in that left column, top to bottom. The order on screen is the order you should fill it in, because the model you pick decides which controls appear underneath.
Step 1: Choose the generation type
Open AI FILMS Studio. The Select Generation Type dropdown holds five entries. Image to Video, Text to Video, Start-End Frame Video, Draw to Video and Video Enhancer.
Choose Text to Video. Changing the type rebuilds the form beneath it rather than navigating to a new page, so switching between modes keeps your place.
Text to Video is the only entry here that needs no input file at all.
Five generation types. Text to Video is the second entry.
Step 2: Select Seedance 2.5
The Select Model dropdown prints the resolution ceiling and the duration range beside every name. Seedance 2.5 sits at the top with 480p-4K · 4s-30s.
The model list. FLUX 3 stops at 1080p and 20 seconds, MiniMax H3 at 2K and 15 seconds.
Those two numbers are limits rather than suggestions. Picking a model with a 1080p ceiling removes the 4K option from the resolution control further down.
Read the list before you commit. If the delivery needs 4K or a clip longer than 20 seconds, Seedance 2.5 is the only model here that can produce it.
Step 3: Write the prompt
The Detailed Prompt field accepts 2,500 characters and shows a live counter in the corner. Its placeholder asks you to describe in detail what you want in your video.
Describe the movement, not only the scene. A prompt that names a subject and stops leaves the camera to the model, and the model will invent one.
Name the shot size, the direction of travel, and what stays still. A 30 second clip needs more of this than a 5 second one, because the model has to decide what happens in the middle.
The prompt field. The counter in the corner reads 0 / 2500.
Step 4: Pick the aspect ratio
Seven tiles are offered, each drawn as the frame shape it produces. 16:9 landscape, 4:3 standard, 1:1 square, 3:4 portrait, 9:16 portrait, 21:9 ultra wide, and an adaptive option that lets the model choose.
Seven aspect ratios. The seventh tile, adaptive, hands the decision to the model.
Aspect ratio is a composition decision taken before generation. The model stages a subject differently in a vertical frame than in a widescreen one, so a vertical cut and a widescreen cut of the same idea are two generations.
The adaptive tile is worth using when you have no fixed delivery format and want the model to pick what suits the described shot.
Step 5: Set the resolution
Video Resolution is a row of four radio buttons. 480p, 720p, 1080p and 4K, with 720p selected by default.
Resolution and duration together drive both the wait and the credit cost, and they multiply. A 30 second clip at 4K is the most expensive combination available in the video workspace.
Test at 480p. Confirm the motion reads correctly, then raise the resolution once the prompt is settled.
Four resolutions. 720p is selected when the form loads.
Step 6: Set the duration
Duration is a continuous slider running from 4 to 30 seconds, marked at 4s, 10s, 20s and 30s. It starts at the low end.
A continuous slider, so any length between 4 and 30 seconds is available.
This is a slider rather than a set of fixed choices, which the older Seedance variants use. Seven seconds is available here and is not available on a model offering only 5, 10 and 15.
Cost scales directly with the number you set. Leave it at 4 seconds while you are still testing the prompt.
Step 7: Decide on audio
Generate Audio is a single toggle, on by default. Seedance 2.5 produces the soundtrack in the same pass as the picture, so dialogue, effects and room tone arrive aligned to the action.
Mention the sound you want inside the prompt. "Rain on a metal roof", "a shutter clicks as the subject turns", and similar cues are read alongside the visual description.
Turning the toggle off lowers the credit cost. The exact figure updates above the Create button as soon as you change it.
One toggle. Audio is generated with the picture rather than added afterwards.
Step 8: Add reference images
Reference Images is marked "Optional, up to 30". Five sources are available. Upload, Previous Task, Image URL, AI Actor and AI Character.
Up to 30 reference images. JPEG, PNG, WEBP, GIF and AVIF, to 300MB, minimum 300px by 300px.
Reference images anchor style, character appearance and scene composition. The AI Actor and AI Character sources pull a face you already built in the Studio, which keeps a person consistent across separate generations.
Previous Task is the fastest route when the anchor is something you generated an hour ago. It saves downloading a file and uploading it again.
Step 9: Add reference videos
Reference Videos is marked "Optional, up to 10" and accepts MP4 only, to 1000MB per file. The same three sources apply. Upload, Previous Task and Video URL.
A reference video supplies motion rather than appearance. Use it when the camera move or the action already exists in footage and you want the model to follow it.
Attaching reference videos changes how the generation is priced, because the length of the reference material counts toward the billed time. The Credits required line updates as soon as you add one, so read it before you press Create.
Up to 10 reference videos. MP4 only, to 1000MB each.
Step 10: Add reference audio
Reference Audios is marked "Optional, up to 10". It accepts MP3 and WAV, to 20MB and a maximum of 30 seconds per file.
Up to 10 reference audio files. MP3 and WAV, to 20MB and 30 seconds each.
Reference audio guides the rhythm and character of the generated track. It is the section that makes music driven work practical, because the cut can follow a beat you supplied.
The three reference sections add up to the 50 asset capacity announced for the model. 30 images, 10 videos and 10 audio files.
Press Create when the form is filled. The Credits required line above the button shows the exact cost for the settings you chose, and the clip appears in the player on the right when it finishes.
Text to Video in the Nodes Graph Editor
The same model is available as a node in the Nodes Graph Editor, which is where Seedance 2.5 becomes part of a longer pipeline rather than a single request.
Connect a Prompt node to a Text to Video node, then wire the output to a Video Viewer node. The Text to Video node shows a green "Prompt connected" line once the wire is in place.
An amber line reads "Connect a negative prompt node (optional)". A negative prompt is a second Prompt node listing what to keep out of the clip, and it is where exclusions belong.
The node reads ~4500 credits for 1080p at 5 seconds. The side panel carries the same four controls as the workspace form.
Selecting the node opens a side panel holding Aspect Ratio, Resolution, Video Duration and Generate Audio. The settings are the same ones the workspace form offers, arranged for a graph rather than a column.
The finished clip opens full width in the viewer, with Download and Delete beneath it.
Click the output thumbnail on the node to open the clip full width. Download saves the file, Delete removes it from the graph.
The node keeps its result after it runs, so you can change one upstream setting and compare the new output against the previous one without rebuilding the graph.
Image to Video
Image to Video animates a still. Seedance 2.5 preserves the subject, composition and lighting of your source frame, then applies the movement your prompt describes.
The form is shorter than the text-to-video one. There is no aspect ratio control, because the source image sets the frame shape.
The three reference sections are absent too. In this mode the input image and the optional last frame are the reference material.
The image to video form. Credits required reads 4500 for 1080p at 5 seconds with audio on.
Step 1: Choose the generation type
Select Image to Video from the same dropdown. It is the first entry in the list.
Image to Video is the first of the five generation types.
The form rebuilds immediately. An Image Source section replaces the aspect ratio tiles, and a Last Frame Image section appears below it.
Your prompt text survives the switch, so an idea written for text-to-video can be reused against a still without retyping.
Step 2: Select Seedance 2.5
The model list here carries the same range, 480p-4k · 4s-30s. Vidu Q3 Drama is the only other model in this mode that reaches 30 seconds, and it stops at 1080p.
Seedance 2.5 is the only entry in image-to-video that combines a 4K ceiling with a 30 second range.
Everything below this control is filtered by it. Choosing a different model changes which resolutions and durations remain available.
The image to video model list, with Seedance 2.5 at the top.
Step 3: Supply the source image
Image Source offers five routes. Upload Image, Previous Task, AI Actor, AI Character and Image URL. Uploads accept JPEG, PNG, WEBP, GIF and AVIF to 100MB, with a minimum of 300px by 300px.
Five ways in. Upload, a previous generation, an AI Actor, an AI Character, or a URL.
This image sets the aspect ratio of the finished clip. Frame it in the shape you want the video to be, because there is no ratio control to correct it later.
A clean source helps. Clear subject separation, even lighting and few compression artefacts give the model more to track when it builds the motion.
Step 4: Add a last frame, if the ending matters
Last Frame Image is optional and takes the same five sources. Supplying one pins where the clip ends, and the model works out the movement between the two frames.
This is the control to use for a continuation. Take the last frame of a finished clip, make it the first frame of the next one, and the two cut together cleanly.
Keep the two frames similar in subject position, lighting and shape. A large gap between them produces an abrupt transition in the middle of the clip.
Optional. Setting it fixes the composition the clip ends on.
Step 5: Write the prompt
The Detailed Prompt field is the same 2,500 character field, with the same placeholder. What you write in it should change.
The same prompt field, 2,500 characters, in both modes.
The model already has the appearance from your image. Describing the subject again wastes characters that should go on the movement.
Write what changes over the clip. The direction the camera travels, what the subject does, how the light shifts, and what holds still.
Step 6: Set resolution, duration and audio
The last three controls match text-to-video exactly. Four resolutions with 720p preselected, a 4 to 30 second slider, and the Generate Audio toggle.
Resolution here is worth matching to the source. Animating a small image at 4K adds cost without adding detail the source never had.
Run the first pass at 480p and 4 seconds. That combination is cheap enough to test whether the still animates well at all.
The same four resolutions as text to video.
4 to 30 seconds, the same continuous slider.
A still that animates well for 5 seconds does not always hold for 30. Motion invented from a single frame drifts further the longer it runs.
Build up in steps. Confirm at 4 seconds, then 10, before committing credits to a full 30 second pass.
Generate Audio behaves the same way in this mode. The soundtrack is produced with the picture and follows what the prompt describes.
Press Create. The Credits required line above the button reads the exact cost, and the clip replaces the placeholder in the player when it finishes.
The audio toggle, on by default in both modes.
Image to Video in the Nodes Graph Editor
The image-to-video route takes one more node than text-to-video, because the source frame needs its own input.
Image Upload and Prompt both feed the Image to Video node. This run reads ~3600 credits.
Wire an Image Upload node and a Prompt node into an Image to Video node, then send its output to a Result node. The generation node confirms both inputs with green "Image connected" and "Prompt connected" lines.
The Result node collects batch output, so a graph producing four variants gathers them in one place rather than four.
The output opens in the same full width viewer as the text-to-video graph, with Download, Delete and Close beneath it.
Reusing one Image Upload node across several generation nodes is the point of building this in a graph. One source frame, several models or several settings, side by side.
The finished clip, with Download, Delete and Close below the player.
Credit Costs
Cost is driven by resolution and duration, and it is the same in both generation types. The table below covers Generate Audio switched on, which is the default.
| Resolution | 5 seconds | 10 seconds | 30 seconds |
|---|---|---|---|
| 480p | 900 credits | 1,800 credits | 5,400 credits |
| 720p | 1,800 credits | 3,600 credits | 10,800 credits |
| 1080p | 4,500 credits | 9,000 credits | 27,000 credits |
| 4K | 9,000 credits | 18,000 credits | 54,000 credits |
Turning Generate Audio off lowers the figure. Attaching reference videos raises it, because the length of the reference material counts toward the billed time. In both cases the Credits required line above the Create button shows the number for your exact settings, so read it rather than estimating.
Failed generations are refunded automatically. Subscription and credit details are on the AI FILMS Studio pricing page.
The jump from 1080p to 4K doubles the cost, and the jump from 5 to 30 seconds multiplies it by six. A 30 second 4K clip costs 60 times a 5 second 480p test of the same prompt, which is the argument for testing low and finishing high.
For cheaper iteration inside the same family, the Seedance 2.0 Mini tutorial covers the compact model that spans 480p to 4K at 4 to 15 seconds. When the delivery is fixed at full HD, the Seedance 2.0 1080 VIP tutorial covers the variant that never drops below 1080p.
Prompt Tips for Seedance 2.5
Write for the length you set. A prompt that fills 5 seconds leaves 25 seconds unaccounted for. For a long clip, describe a sequence of beats in order rather than a single moment.
Name the camera explicitly. "A slow crane rising above the treeline over the full clip" produces a different result from "the camera moves up". Seedance 2.5 follows cinematographic instruction closely enough that the specific wording matters.
Say what holds still. Video models tend toward constant motion. A locked camera or a stationary subject generally has to be stated, or the model will add movement you did not ask for.
Put exclusions in a negative prompt. In the Nodes Graph Editor that is a second Prompt node wired to the negative input. Telling the main prompt to avoid something tends to produce it.
Describe sound. Audio is generated from the same prompt, so naming the sounds gets them synchronised to the action. Silence is also a choice, and the toggle is cheaper.
In image-to-video, describe only the change. The appearance is already fixed by your source frame. Characters spent redescribing it are characters not spent on the motion.
For a side by side comparison of how the earlier family handles the same two modes, the Seedance 2.0 tutorial walks through both at the 15 second ceiling. The other model in the roster that reaches 30 seconds is covered in the Vidu Q3 Drama and Turbo tutorial, which stops at 1080p and takes a different approach to long takes.
Limitations
The prompt field stops at 2,500 characters. For a 30 second clip that is roughly 80 characters per second of output, so a detailed sequence needs tight writing.
Image-to-video has no aspect ratio control. The source image decides the frame shape. Reframing means preparing a new source image at the ratio you want.
Reference audio is capped at 30 seconds per file. A longer track has to be trimmed before upload, and only the section you supply guides the generation.
4K at 30 seconds is the slowest option in the workspace. It is also the most expensive single generation available, so it belongs at the end of an iteration cycle rather than the start.
Motion drifts on long generations from a single still. Image-to-video at 30 seconds asks the model to invent a great deal from one frame. A last frame image constrains it, and shorter clips joined at matching frames often read better.
Sources
ByteDance Seed | Volcano Engine
Continue Reading
Video & LipSync
- Video Generator
- Text to Video
- Image to Video
- Start-End Frame to Video
- Draw to Video
- Motion Control
- Video Enhancer
- Video Upscaler
- Video to Video LipSync
- Audio to Video LipSync
- Image to Video LipSync
- Video FaceSwap
- Seedance 2.5
- Seedance 2
- FLUX 3
- Minimax H3
- Vidu Q3 Pro
- Grok Imagine Video 1.5
- Gemini Omni
- Google Veo 3.1
- Kling 3.0 Pro
- Luma Ray 3.2
- LTX 2.3
- Happy Horse 1.1
- WAN 2.7
- Kling 3.0 Motion
- ByteDance Upscaler
- InfiniteTalk
- InsightFace


.jpg?w=3840)