EditorNodesPricingBlog

LTX-2.5: Lightricks Ships Open Weights World Model with Native Multishot and 6.8-Second Video Generation

September 1, 2026
LTX-2.5: Lightricks Ships Open Weights World Model with Native Multishot and 6.8-Second Video Generation

Share this post:

LTX-2.5: Lightricks Ships Open Weights World Model with Native Multishot and 6.8-Second Video Generation

Lightricks released LTX-2.5 on August 11, 2026. The model holds character identity, environment, lighting, and voice consistent across multiple shots in a single generation pass. It is the first model in the LTX series to deliver this capability natively, without requiring separate consistency tooling or multiple inference runs.

The weights are released on Hugging Face under the Lightricks organisation. Commercial use is free for companies with annual revenue under $10 million.

LTX-2.5 open weights world model for multishot video generation from Lightricks

Native Multishot Consistency

Every open source video model released before LTX-2.5 produced single shot output. A single shot is a continuous clip with one camera setup. A narrative scene, by contrast, is a sequence of shots: wide, medium, close up, cutaway, back to wide. Producing a coherent scene previously required running multiple generation passes and applying consistency tooling between them. Character drift, lighting inconsistency, and voice mismatch across cuts were the result.

LTX-2.3, released in March 2026, added native audio generation, making it the first model in the series to produce video with synchronized sound in a single pass. It did not address multishot consistency.

LTX-2.5 addresses both. A single generation pass produces a connected sequence of shots with the same character appearance, the same environment, the same lighting conditions, and the same voice profile across every cut. The model treats the full scene as one generation problem rather than treating each shot as a separate problem.

For AI filmmakers, this changes the production unit from the shot to the scene. Prior workflows required assembling scenes from independently generated shots with post processing to reduce inconsistency. LTX-2.5 produces the scene as a unit.

Architecture: 22B Diffusion Transformer

LTX-2.5 uses a 22 billion parameter Diffusion Transformer architecture. The model adds two components not present in earlier LTX releases.

The first is a new diffusion video decoder. Earlier models used standard decoders shared with image generation architectures. The new decoder is designed specifically for video, handling temporal coherence across frames and across cuts in a way that image decoders cannot address.

The second is a Gemma 4 12B text encoder. The text encoder is responsible for translating prompt language into the conditioning signal that guides generation. A 12B parameter encoder provides substantially more language understanding than the smaller encoders used in previous open source video models. This matters for following complex multishot prompts that describe scene transitions, character actions across cuts, and changes in environment within a single generation.

The combination of the diffusion video decoder and the Gemma 4 12B encoder is the architectural basis for the multishot consistency capability. Each component addresses a different failure mode in multishot generation.

Speed Benchmark

LTX-2.5 generates a 10-second 720p image-to-video clip in 6.8 seconds on two NVIDIA GB200 GPUs. That benchmark places it among the fastest open weights video generation models available at the 720p resolution.

The GB200 hardware is Nvidia's current high end infrastructure class GPU. The benchmark was run on two units, which is an infrastructure configuration rather than a consumer setup. For teams running on consumer hardware, generation times will be longer. The benchmark establishes the performance ceiling under current conditions.

Output resolution ranges from 720p to native 4K. Clip length ranges from 6 to 20 seconds. For a 20-second clip at 4K, generation time on GB200 hardware exceeds the 6.8-second benchmark. The benchmark specifically covers a 10-second 720p image-to-video generation case.

Resolutions and Clip Lengths

The model generates at 720p, 1080p, and native 4K. The 4K output is native, not upscaled from a lower resolution. This is a different approach from earlier models in the LTX series, which generated at lower resolutions and applied a separate upscaler to reach 4K.

The upscaler approach was part of the LTX-2.3 workflow, where external upscalers were required for 4K output. LTX-2.5 eliminates that step by building 4K generation into the base model.

Clip length from 6 to 20 seconds is a wider range than previous LTX releases. For scene level generation, 20 seconds provides enough duration for a three or four shot sequence with meaningful action in each shot. Most narrative scenes in film and television run between 15 and 90 seconds. A single 20-second generation covers a substantial opening or closing beat within that range.

Licensing and Commercial Use

LTX-2.5 is released as open weights on Hugging Face. Commercial use is free for companies with annual revenue under $10 million. Companies above that threshold can access the model through the LTX API, which provides managed infrastructure rather than requiring direct weight deployment.

The revenue threshold is a specific commercial model that Lightricks introduced with this release. It separates the hobbyist and small business use case, which gets free access to weights, from the commercial infrastructure use case, which gets managed API access.

Open weights means the parameters are publicly available for download, modification, and local deployment. This is different from a model accessible only through an API. A team running LTX-2.5 on their own hardware has full control over generation without sending data to Lightricks' infrastructure.

The Apache style license terms make the open weights approach available for commercial integration in products and workflows, provided the revenue threshold is met or API access is used. This positions LTX-2.5 for integration into production tools, fine tuning pipelines, and custom deployment configurations.

Fine-Tuning with LTX Trainer

LTX Trainer, released in June 2026, provides a unified fine tuning framework for the LTX model family. It supports LoRA fine tuning, scene fine tuning, and character fine tuning within a single framework. LTX-2.5 is compatible with LTX Trainer, making it possible to fine-tune the 22B model on custom character assets, environment references, or style references.

For production teams building with specific characters or settings across multiple scenes, fine tuning LTX-2.5 on reference material provides a consistent generation baseline. The multishot consistency capability of the base model is then combined with character specific fine tuning to produce scene level output that holds both to the model's native consistency and to the custom reference material.

Fine Tuning a 22B model requires substantial compute. LTX Trainer supports efficient fine tuning approaches that reduce the compute requirement relative to full fine tuning. The specific requirements depend on the fine tuning configuration, the training data, and the hardware available.

Open Weights in the Context of Prior LTX Releases

LTX-2 launched in January 2026 as an open source 4K video model. It established Lightricks as a serious participant in the open weights video generation space. The model generated single shot output and was the first in the series to produce 4K resolution video as an open weights release.

LTX-2.3 followed in March 2026 with native audio generation, extending the single shot model to include synchronized sound. LTX Trainer arrived in June 2026 to support fine tuning across the model family.

LTX-2.5 is the fourth release in this arc and the most architecturally significant. The jump from single shot to multishot generation is a larger capability step than the jump from silent to audio. Multishot consistency is the specific capability that separates a video generation model from a filmmaking tool. LTX-2.5 is the first open weights model to make that claim clearly.

The 22B parameter count is the largest in the LTX series. The increase in model scale is directly connected to the multishot capability. Holding character, environment, and voice consistent across cuts requires representing more context than single shot generation, and that additional representation capacity requires additional parameters.

Consumer Hardware Deployment

The GB200 benchmark is the upper bound. Deployment on consumer hardware is the practical baseline for most teams working with the model.

RunPod's deployment guide confirms that LTX-2.5 runs on consumer GPU configurations through cloud instances. The minimum viable configuration for 720p generation requires a card with at least 24GB VRAM. For 4K generation, more VRAM is required. The exact minimum depends on the clip length and the number of shots in the generation.

The practical implication for smaller teams is that 720p multishot generation is accessible on current consumer hardware, while 4K generation at the model's full output range requires more infrastructure. The 6.8-second benchmark on GB200 does not translate directly to consumer hardware performance. Generation times on consumer configurations are longer, but the same multishot capability is present regardless of hardware.

For teams prioritizing speed over resolution, 720p multishot generation at consumer hardware performance times is a usable production workflow. For teams prioritizing resolution, the API route gives access to faster infrastructure without managing weights deployment directly.

The LTX Model Arc

LTX-2.5 is the fourth release in a model arc that began in January 2026. The arc shows a consistent progression toward capabilities that filmmakers need.

LTX-2 delivered 4K resolution as an open weights release, establishing that 4K output did not require proprietary infrastructure. LTX-2.3 added native audio, eliminating the separate voice generation step in single shot workflows. LTX Trainer added fine tuning capability, making character and environment customization accessible without full model retraining. LTX-2.5 adds multishot consistency and the new diffusion video decoder, completing the basic capability set for scene level generation.

The arc from LTX-2 to LTX-2.5 covers the eight months from January to August 2026. The pace of development across those releases has been faster than typical model development cycles. Each release added a capability that filmmakers had identified as a specific gap in the prior model.

What Multishot Consistency Enables

A scene consists of shots edited together. The editing creates the narrative. What makes a scene feel coherent is consistency across those shots: the same actor looks the same in the wide and the close up, the room has the same furniture in the cutaway as in the establishing shot, the voice that spoke in the first shot is recognizably the same voice in the fourth.

AI video generation before LTX-2.5 could not produce this reliably from a single prompt. Each shot had to be generated separately, with drift between them managed through post processing. The result was sequences that looked assembled rather than captured.

LTX-2.5 changes what is possible in a single pass. The model generates the sequence with the consistency decisions already made internally. The output is a scene, not a collection of shots.

For AI filmmakers, this affects the creative process at the planning stage. A scene can be described as a prompt, generated as a unit, and evaluated as a whole. Adjustments are made to the prompt or to fine tuning parameters, not to an assembly of inconsistent parts.

The specific prompt engineering for multishot generation is different from single shot prompting. A multishot prompt must describe the sequence of shots, the transitions between them, and the action within each shot in a way that the model can interpret as a connected temporal sequence. The Gemma 4 12B text encoder handles more complex prompt structure than smaller encoders, which is part of why the encoder choice is architecturally significant for this capability.

Teams working with LTX-2.5 for the first time will need to develop prompting approaches suited to the model's multishot architecture. Single shot prompting habits, which describe one continuous clip, translate poorly to multishot generation. A multishot prompt reads more like a scene description with shot breakdowns than a single action description.

The open weights release makes that experimentation freely available. Teams can iterate locally on prompt approaches without incurring inference costs per generation. For a model with a capability as new as native multishot generation, local iteration access is valuable. The community's prompting patterns for multishot generation will develop quickly once teams have access to the weights.

The 22B parameter count also means that fine tuning on custom character assets produces stronger results than fine tuning on smaller models. A larger base model holds more representation capacity, which means fine tuning on a specific character does not have to overwrite as much base model capability to produce consistent character output. The LTX Trainer framework makes this fine tuning pathway directly accessible without requiring custom training infrastructure.

LTX-2.5 is available on Hugging Face under the Lightricks organisation as of August 11, 2026. The model card includes generation examples, prompt guidance, and hardware requirements for different resolution targets. Teams evaluating the model for production workflows can start with the published examples before moving to custom prompt development.

The model represents the current state of open weights video generation as of August 2026. The multishot capability is the defining addition. The 22B scale, the new diffusion video decoder, and the Gemma 4 12B text encoder are the architectural choices that make that capability possible at the resolution and speed ranges Lightricks has published.

Lightricks has not announced what follows LTX-2.5 or a timeline for the next release. The August 2026 release concludes the arc that began with LTX-2 in January.

Filmmakers who want to work with AI video generation tools can access LTX-2.5 and other open weights models through AI FILMS Studio's video workspace.


Sources

VentureBeat | DataNorth AI | Open Source For You | LLM Stats | RunPod Blog