EditorNodesPricingBlog

Grok Imagine Video 1.5 Tutorial: Text to Video and Image to Video

August 8, 2026
Grok Imagine Video 1.5 Tutorial: Text to Video and Image to Video

Share this post:

Grok Imagine Video 1.5 Tutorial: Text to Video and Image to Video

Grok Imagine Video 1.5 from xAI is now available on AI FILMS Studio. The model generates video from a text prompt or from an uploaded still, at any whole second length from 1 to 15. That 1 second floor is the lowest in the entire video roster. No other model in the workspace will render a clip shorter than 2 seconds.

The floor matters because it changes what a test costs. A 1 second draft at 480p bills 80 credits, so you can check composition, motion direction and framing for less than the price of a coffee stirrer before committing to the full length.

This tutorial covers both generation modes step by step, plus the node setup for each one in the Nodes Graph Editor. The model sits in both the text to video generator and the image to video generator. It is the same family behind the Grok Odyssey feature film announcement, now available to run yourself.

What Is Grok Imagine Video 1.5

Grok Imagine Video 1.5 is xAI's video generation model. In text-to-video mode it builds a clip from a written description alone. In image-to-video mode it takes your uploaded frame as the opening shot and animates forward from it.

The two modes carry different names in the workspace. The text-to-video dropdown lists it as Grok Imagine Video 1.5. The image-to-video dropdown lists it as Grok Image Video 1.5. Same model, two labels. Scan for "Grok" rather than for "Imagine" when you switch modes, or you will scroll past it.

Video resolution. Text-to-video offers 480p, 720p and 1080p. Image-to-video offers 480p and 720p only, with no 1080p option. Both default to 720p. This is the one place where the two modes genuinely differ in output quality, so plan a 1080p delivery around text-to-video.

Duration. A slider from 1 to 15 seconds in whole seconds, identical in both modes, defaulting to 6. Cost tracks duration directly. The slider is the control that decides most of what you pay.

Aspect ratios, text-to-video only. Five formats: 16:9 Landscape, 9:16 Portrait, 1:1 Square, 3:2 Widescreen and 2:3 Portrait. 16:9 is selected by default. Image-to-video has no aspect ratio control at all, because the source image sets the output shape.

Prompt length. The Detailed Prompt field accepts up to 2,500 characters in both modes. That covers camera behaviour, subject action, lighting and environment in one pass without running out of room.

No negative prompt, no seed. Neither control appears in the workspace for this model. Exclusions have to be written into the positive prompt. The Nodes Graph Editor does accept an optional negative prompt node, which is the only route to one.

Text to Video

Step 1: Open the video workspace

Go to AI FILMS Studio. The Video Generator panel loads on the left and generated clips appear on the right.

Grok Imagine Video 1.5 text to video generation workspace interface on AI FILMS Studio

Step 2: Select Text to Video as the generation type

Open the Select Generation Type dropdown at the top of the panel and choose Text to Video. The same dropdown carries Image to Video, Start-End Frame Video, Draw to Video and Video Enhancer.

Select Generation Type dropdown showing Text to Video in the AI FILMS Studio video workspace

Step 3: Select Grok Imagine Video 1.5

Open the Select Model dropdown and choose Grok Imagine Video 1.5. Every row shows the resolution ceiling and the duration range for that model, so the Grok row reads 480p to 1080p and 1s to 15s. Comparing that row against the ones above and below it is the fastest way to rule a model in or out.

Model dropdown showing Grok Imagine Video 1.5 at 480p to 1080p and 1 to 15 seconds on AI FILMS Studio

Step 4: Write your prompt

Enter your description in the Detailed Prompt field. Lead with the camera instruction, then the subject, then the environment and lighting. The counter under the field tracks your usage against the 2,500 character limit.

There is no negative prompt field here, so anything you want kept out of frame has to be phrased positively. Write "an empty road at dawn" rather than relying on a negative to remove traffic.

Detailed Prompt field with a 2500 character counter for Grok Imagine Video 1.5 on AI FILMS Studio

Step 5: Choose aspect ratio

Pick one of the five tiles. 16:9 Landscape suits cinema and desktop delivery. 9:16 and 2:3 Portrait suit vertical feeds. 1:1 Square covers social placements that crop badly at wider ratios, and 3:2 Widescreen sits between 16:9 and 4:3 for stills work carried into motion.

Set this before generating. The ratio cannot be changed afterwards without a new generation.

Grok Imagine Video 1.5 aspect ratio tiles showing 16:9, 9:16, 1:1, 3:2 and 2:3 options

Step 6: Choose the resolution

The Video Resolution control offers 480p, 720p and 1080p, with 720p selected by default. The three tiers are the reason to iterate here rather than elsewhere. A 480p pass costs a third of a 1080p pass at the same length.

Use 480p while you are still testing the prompt. Move to 720p once the shot works, and reserve 1080p for the take you intend to deliver.

Video Resolution control for Grok Imagine Video 1.5 showing 480p, 720p and 1080p options

Step 7: Set the duration

Drag the Duration slider to any whole second between 1 and 15. The slider carries marks at 1s, 5s, 6s, 10s and 15s, and it opens at 6.

Pull it all the way down to 1 second for the first pass on a new prompt. One second is enough to see whether the model understood the subject, the framing and the direction of movement. Everything beyond that is the same shot continuing.

Duration slider for Grok Imagine Video 1.5 running from 1 second to 15 seconds

Step 8: Generate

Check the credit estimate, then click Generate. The clip appears in the output panel on the right when processing finishes, with Download and Delete controls beneath it.

Text to Video in the Nodes Graph Editor

Grok Imagine Video 1.5 text-to-video also runs as a node in the Nodes Graph Editor. Add a Prompt node and connect it to a Text to Video node. Open the model dropdown inside the node, choose Grok Imagine Video 1.5, then set resolution, duration and aspect ratio in the node settings panel.

Grok Imagine Video 1.5 text to video node workflow connecting Prompt, Text to Video and Video Viewer

The node card shows its credit cost before you run it, so a 10 second generation at 720p reads roughly 1400 credits. The card also confirms the prompt connection and offers an optional negative prompt node. That node is the only negative prompt this model has, and it is worth wiring when a specific element keeps appearing in output you did not ask for.

Connect the output to a Video Viewer node to play the result inside the graph. From there you can download the clip or pass it into another node, such as a video enhancer or a lipsync stage.

Grok Imagine Video 1.5 text to video result playing in the Video Viewer node of the Nodes Graph Editor

Image to Video

Image-to-video animates a still into a clip. The model reads your uploaded image as the first frame and carries its subject, lighting and composition forward through the generation.

Step 1: Switch the generation type to Image to Video

Open the Select Generation Type dropdown and choose Image to Video. The panel reloads with an Image Source section above the prompt field.

Grok Image Video 1.5 image to video interface with the Image Source section on AI FILMS Studio

Step 2: Select Grok Image Video 1.5

Open the Select Model dropdown and choose Grok Image Video 1.5. The row reads 480p to 720p and 1s to 15s. The missing 1080p tier is the difference from text-to-video, and the dropdown states it before you commit.

Model dropdown showing Grok Image Video 1.5 at 480p to 720p for image to video on AI FILMS Studio

Step 3: Supply the source image

The Image Source section offers five inputs. Upload Image takes a file from your machine. Previous Task pulls a frame from something you already generated. AI Actor and AI Character pull from your saved characters. Image URL takes a link.

Uploads accept JPEG, PNG, WEBP, GIF and AVIF up to 100MB, at a minimum of 300px by 300px.

Image Source section with Upload Image, Previous Task, AI Actor, AI Character and Image URL options

Step 4: Write your motion prompt

Describe what moves and how the camera behaves. Skip the appearance of the subject, because the model already reads that from your uploaded frame. "The camera pushes in slowly as rain streaks the foreground" gives the model more to work with than a restatement of what is already in the picture.

The 2,500 character limit is the same as text-to-video, and there is still no negative prompt field.

Detailed Prompt field for Grok Image Video 1.5 image to video generation on AI FILMS Studio

Step 5: Choose the resolution

Two options here, 480p and 720p, with 720p selected by default. 1080p is absent in this mode.

Frame and crop your source image to the delivery format before you upload it. There is no aspect ratio control in image-to-video, so whatever shape goes in is the shape that comes back. Fixing the ratio afterwards means regenerating from a recropped image.

Video Resolution control for Grok Image Video 1.5 showing 480p and 720p options

Step 6: Set the duration

The Duration slider runs 1 to 15 seconds in whole seconds, the same range as text-to-video, with marks at 1s, 6s, 10s and 15s.

A 1 second pass is useful here for a different reason. It shows you whether the model respects your source frame before you spend on a longer clip. Identity drift and warping on the subject show up in the first second when they are going to show up at all.

Duration slider for Grok Image Video 1.5 running from 1 second to 15 seconds

Step 7: Generate

Review the credit estimate and click Generate. The animated clip appears in the output panel and can be downloaded from there.

Image to Video in the Nodes Graph Editor

The image-to-video mode runs as a node too. Add a Prompt node and an Image Upload node, and connect both into an Image to Video node. Select Grok Image Video 1.5 in the node model dropdown, then set resolution and duration in the node settings panel.

Grok Image Video 1.5 image to video node workflow with Prompt, Image Upload and Image to Video nodes

The node confirms both connections before it will run, showing Image connected and Prompt connected on the card. A 5 second generation at 720p reads roughly 710 credits on the card. As with text-to-video, a negative prompt node is optional.

Connect the output to a Video Viewer node to review the clip inside the graph. This layout is worth saving when you animate a batch of stills, because only the Image Upload node changes between runs.

Grok Image Video 1.5 image to video result playing in the Video Viewer node of the Nodes Graph Editor

Credit Costs

Text-to-video bills duration multiplied by a rate that depends on the resolution. Image-to-video uses the same per second rates and adds a flat 10 credit charge for processing the source image.

Text to Video

Duration 480p 720p 1080p
1s 80 140 250
5s 400 700 1,250
6s 480 840 1,500
10s 800 1,400 2,500
15s 1,200 2,100 3,750

Image to Video

Duration 480p 720p
1s 90 150
5s 410 710
6s 490 850
10s 810 1,410
15s 1,210 2,110

Cost moves in two dimensions at once here, because both resolution and duration change the price. A 15 second clip at 1080p costs 46 times a 1 second clip at 480p. The MiniMax H3 tutorial covers the opposite arrangement, a model fixed at 2K with a 5 second floor, where duration is the only variable in play.

Credits for failed generations are refunded automatically. For subscription details, see AI FILMS Studio pricing.

Prompt Tips for Grok Imagine Video 1.5

Test at 1 second and 480p. An 80 credit test is the cheapest iteration loop available in the video workspace. Run three or four of them on a new prompt before you touch the resolution control. The first second tells you whether the model read the subject and the camera direction correctly, which is most of what goes wrong.

Climb the resolution tiers in order. 480p to confirm the idea, 720p to confirm the detail, 1080p for the delivery take. Going straight to 1080p on an untested prompt costs three times as much per second and tells you nothing extra about composition.

Write exclusions as inclusions. With no negative prompt field, the only way to keep something out of frame is to fill that space with something else. "A bare concrete wall behind the subject" works. Wishing for no text on the wall does not.

Crop the source image before uploading. Image-to-video has no aspect ratio control. A 16:9 photograph produces a 16:9 clip, whatever the target platform wants. Crop to 9:16 first when the clip is going to a vertical feed.

Lead with the camera move. Put the camera instruction at the front of the prompt. "Slow crane down over a rain soaked street as two figures cross the frame" reads better than the same scene with the camera note added at the end.

Use the graph for negative prompts. When a specific artefact keeps appearing, rebuild the shot in the Nodes Graph Editor and attach a negative prompt node. The workspace form cannot do this, and it is the one capability the graph adds for this model.

For a model that holds a fixed 2K at every length, the MiniMax H3 tutorial covers the trade in the other direction. The LTX 2.3 tutorial covers a faster model with frame rate control, the Vidu Q3 Drama and Turbo tutorial covers script driven sequences across multiple reference assets, and the Seedance 2.0 tutorial covers camera movement and scene coherence.

Sources

xAI: xAI