Back to Model List

LTX-2.5 – LTX's Open-Source AI Video Generation Foundation Model

AI Tech Editorial
RSS Feed
LTX-2.5 – LTX's Open-Source AI Video Generation Foundation Model official screenshot
(Image source: official screenshot)

Executive Summary:

LTX-2.5 is a professional-grade AI video generation foundation model open-sourced by Lightricks, an Israeli AI company. With 220 billion parameters, it provides full model weights upon release. The mo...

1. What is LTX-2.5

LTX-2.5 is a professional-grade AI video generation foundation model open-sourced by Lightricks, an Israeli AI company. With 220 billion parameters, it provides full model weights upon release. The model introduces multiple architectural innovations, supporting native multi-shot continuous generation, 4K HDR, and RAW/EXR lossless format outputs, and can simultaneously generate matching audio tracks. Through precise video editing and various reference control modes (Canny edge, Depth, Pose), LTX-2.5 achieves significant breakthroughs in narrative coherence and creative controllability, offering a high-cost-performance open-source alternative for cinematic-grade AI video production.

ltx-2-5-ltx-ai official website screenshot
Image source: Official article
Image source: official article

Technical positioning and domain: LTX-2.5 is an AI video generation model at the intersection of computer vision and natural language processing, positioned as a high-quality video generation tool for professional film production, advertising creativity, and AI short drama creation. Compared to mainstream closed-source video generation models, it differentiates itself through full open-sourcing, support for local deployment and fine-tuning, filling the gap in the open-source ecosystem for 4K-level professional video generation models.

Development background: This model was developed by the Lightricks team, which has accumulated years of technical expertise in the field of AI image and video editing (products such as Facetune and LTX Studio). The release of LTX-2.5 aims to address the industry trend of increasingly high prices for closed-source models, providing the film industry with controllable, customizable, and auditable professional-grade video generation capabilities through an open-ecosystem strategy, significantly lowering the technical barriers and costs of high-quality AI video creation.

Core value: LTX-2.5 addresses two long-standing pain points in AI video generation—shot continuity and creative controllability. Its native multi-shot generation capability allows for multiple continuous and consistent shots in a single output, fundamentally avoiding the disconnection between characters, scenes, and lighting during transitions. Precise video editing and the IC-LoRA control mode enable creators to shift from "random card drawing" to a precise production workflow of "real-shot skeleton + AI stylized overlay," significantly enhancing usability in professional scenarios.

Technical features: The model employs a two-stage rendering strategy of Diffusion Fidelity Rendering. It first constructs a motion skeleton in a high-compression latent space, then parses texture details with pixel-level precision, achieving adaptive computational resource allocation. It includes a custom Gemma 4 12B text encoder and a Prompt Enhancer auxiliary model, which can fully preserve complex long prompts' multi-subject relationships and action details. An independent Audio VAE module enables synchronized audio-video generation, while the Duration Head module gives the model an initial perception of narrative pacing, automatically inferring the optimal video length.

2. Key Features

  • Native Multi-Camera Generation: A single inference can generate multiple continuous shots with different camera angles and compositions, maintaining high consistency in character appearance, environmental atmosphere, lighting conditions, and sound during shot transitions. This capability fundamentally addresses the core challenge of shot continuity in AI-driven storytelling, allowing AI-generated short films to be narrative segments with editing rhythm, rather than monotonous single shots.

  • Diffusion Fidelity Rendering: Utilizes a two-stage rendering strategy of "structure first, then details." In the first stage, motion, composition, and shot framework are built in a latent space with 8× temporal compression, and high-fidelity keyframes are adaptively generated. In the second stage, the final video is rendered by jointly diffusing from the structure and keyframes, parsing textures, materials, and facial details with pixel-level precision. The system dynamically allocates computational resources based on scene complexity, allocating more resources to complex dynamic shots and automatically saving resources for static scenes, thereby improving overall efficiency while maintaining visual quality.

  • Precise Video Editing: Performs high-precision style rewriting, scene repainting, and character replacement while preserving the original motion trajectory and spatial structure of the live-action footage. This feature enables a controllable creation mode of "live-action skeleton + AI stylized overlay," allowing creators to first shoot action footage using traditional methods and then replace the visual style using AI, greatly enhancing flexibility and control in professional film post-production.

  • IC-LoRA Precise Control: Offers three reference video control modes: Canny edge, Depth, and Pose. Users can input a reference video into the model and use the corresponding control mode to lock composition, spatial structure, or character movement, ensuring the generated output strictly follows the skeleton of the reference material. This enables a transition from "prompt-driven" to "reference-driven" generation, offering significant practical value in advertising and film compositing.

  • Automatic Duration Prediction (Duration Head): Includes a built-in Duration Head prediction module, which is deployed before the diffusion process. This module automatically infers and outputs the most appropriate video duration by understanding the action semantics in the prompts. Creators no longer need to repeatedly test different durations; the model has an initial perception of narrative pacing and can automatically match durations based on descriptions such as "slow-motion running" or "rapid scene switching."

  • Synchronized Audio Generation: Utilizes an independent Audio VAE module to generate synchronized audio tracks for the produced video. This module is an audio latent space encoder that maps video content to corresponding audio representations, enabling synchronized generation of visuals and sound. For short-form video and advertising production scenarios requiring quick generation of background music or ambient sound, this feature significantly reduces the workload of post-production audio effects.

  • Professional-Grade Output Formats: Supports native 4K HDR, RAW, and EXR lossless output formats. These formats can be directly integrated into professional film grading, visual effects compositing, and post-production editing pipelines (such as DaVinci Resolve and After Effects), without the need for compression or transcoding, meeting the industry's rigid demand for lossless source material. For users seeking color depth and post-production flexibility, this is a core capability typically not offered by closed-source models.

3. How to Use

  1. Environment Requirements and Preparation: LTX-2.5 is recommended to be run on hardware with at least a 24GB GPU memory (such as NVIDIA RTX 4090, A100, or higher), and the operating system should be Linux or Windows 10/11. Python 3.10 or higher and CUDA 11.8 or higher drivers must be installed. Local deployment primarily relies on the ComfyUI workflow platform, and it is essential to ensure that ComfyUI is updated to the latest version.

  2. Install the LTXVideo Node: After launching ComfyUI, use the built-in Manager tool to search for and install the "ComfyUI-LTXVideo" node package. This node package includes core functional modules such as model loading, parameter configuration, and multi-shot video generation. Restart ComfyUI after installation to activate the node.

  3. Download the Model and Configure the Workflow: In the ComfyUI template browser, search for "LTX-2.5" and select the official workflow template. The template will automatically link all required model files (including base weights, text encoder, VAE, etc.), and you can download them with one click. Alternatively, users can download the model weights directly from HuggingFace and manually place them in the models directory of ComfyUI.

  4. Parameter Configuration: In the LTXVideo node, configure key parameters: width and height must be multiples of 32 (e.g., 1024×576), the number of frames must be a multiple of 1+8 (i.e., 9, 17, 25, etc.), and the frame rate can be selected from 24, 25, 48, or 50 fps. It is recommended to set the base resolution to half of the target resolution (e.g., set base to 540p if the target is 1080p), and the model will automatically upscale by 2×. The output format can be chosen from HDR, RAW, or EXR, depending on post-production requirements.

  5. Writing Prompts and Executing Generation: Prompts should be written as a coherent paragraph of descriptive text, including shot composition (e.g., medium shot, close-up), scene lighting (e.g., warm golden hour light), character actions (e.g., running, smiling), camera movement (e.g., zoom in, pan), and audio cues (e.g., background wind sound). Avoid using tag stacking or weight syntax. Click "Queue Prompt" to execute the generation: text-to-video will generate the video directly; image-to-video requires uploading the first frame as the starting frame; the start-end frame mode requires uploading both the starting and ending frames simultaneously.

  6. Advanced Control and Model Selection: Advanced users can load IC-LoRA models in Canny, Depth, or Pose modes and integrate reference videos to lock in composition or motion. For model selection, use the distilled model for quick iterations (faster inference speed), and use the full dev model for the final output to achieve more refined details. It is recommended to first test with low resolution and short duration, then increase parameters once the results are confirmed.

4. Pros and Cons Analysis

Pros
Fully Open-Source Weights: The 22B parameter model opens its weights under the Apache 2.0 license, supporting local deployment, commercial use, and fine-tuning. This provides full transparency and customization options for film studios and researchers.
Native Multi-Shot and Audio-Video Synchronization: The model can generate multiple consecutive shots in a single run while maintaining consistency and synchronizing audio. This is a first in open-source video models and significantly improves storytelling efficiency and the completeness of the final output.
Industrial-Grade Output Formats: Native support for 4K HDR, RAW, and EXR lossless formats allows direct integration into professional post-production pipelines without the need for transcoding or compression. This meets the production standards for cinema and high-end advertising, a capability rarely offered by closed-source models.
Adaptive Rendering Efficiency: Diffusion Fidelity Rendering dynamically allocates computational resources based on scene complexity, ensuring pixel-level detail while avoiding unnecessary computation for static scenes. As a result, the overall generation efficiency is higher than that of similar open-source models.

5. Comparative Analysis with Similar Tools

Comparison Dimension LTX-2.5 Minimax H3 CogVideoX-5B
Developer Lightricks (Israel) MiniMax (China) Zhipu AI (China)
Openness 22B weights fully open-sourced, supports local deployment and fine-tuning Partial weight openness (excluding US/UK/EU/Korea), no support for local fine-tuning 5B model weights open-sourced, supports local deployment
Parameter Scale 22B (public) Not disclosed 5B (public)
Maximum Resolution 4K (3840×2160) 2K (fixed) 768×768
Maximum Duration 20 seconds 10 seconds 6 seconds
Native Multi-Camera Support ✅ Supported ❌ Not supported ❌ Not supported
Synchronized Audio ✅ Supported (Audio VAE) ❌ Not open ❌ Not supported
Controllability Precise video editing + IC-LoRA (Canny/Depth/Pose) Primarily prompt-driven Only prompts + limited parameter control
Output Format 4K HDR, RAW, EXR lossless Standard compressed formats MP4

Selection Recommendations: For film studios seeking the highest visual quality, full compatibility with professional post-production pipelines, and maximum creative freedom, LTX-2.5 is currently the top choice among open-source solutions. Its native multi-camera support and synchronized audio capabilities significantly reduce the complexity of short film production workflows, while 4K HDR/RAW output ensures usability of the material in professional color grading pipelines. If the project budget is limited and the resolution requirement is not high (within 1080p), CogVideoX-5B, with its smaller parameter scale and lower hardware demands, can serve as an alternative for rapid prototype validation.

For product teams needing quick deployment into web applications or mobile platforms, Minimax H3's API service offers advantages in latency and usability. However, due to regional restrictions and lack of fine-tuning support, it is unsuitable for enterprises requiring customized models. Runway Gen-3 Alpha maintains stable generation quality and a rich set of features among closed-source models, but its high subscription costs and reliance on online services make it unsuitable for localized, offline professional scenarios. Overall, LTX-2.5 stands out with its openness, output quality, and integration of professional features, making it particularly suitable for teams looking to deeply embed AI video generation into existing film workflows.

6. Editor's Summary

The release of LTX-2.5 marks a new stage of maturity for open-source AI video generation models. From a technological innovation perspective, the model has achieved breakthroughs across multiple dimensions: the two-stage rendering strategy of Diffusion Fidelity Rendering achieves a good balance between computational efficiency and visual detail, avoiding common issues such as motion blur and artifacts in previous open-source models; the native multi-shot generation capability elevates AI video from single-shot segments to narrative-level editing, a feature that the open-source community has achieved for the first time; the introduction of the Audio VAE module makes synchronized audio-visual generation no longer exclusive to closed-source models. These innovations are not merely a stack of features, but systematic solutions addressing specific pain points in professional film production workflows (shot continuity, post-production compatibility, audio-visual synchronization).

In terms of practical value, LTX-2.5's most notable contribution is providing the film industry with a truly usable open-source alternative. The open weights with 22B parameters mean that creators can fully control their data and generation processes, without relying on third-party APIs or worrying about service interruptions or price fluctuations. Its lossless output capability for 4K HDR/RAW/EXR allows it to seamlessly integrate into existing professional post-production pipelines, a function that no previous open-source model has achieved. Additionally, precise video editing and the IC-LoRA control mode transform AI video generation from a "lottery-style" creative process into a "guided" one, significantly enhancing its usability in commercial projects.

This model is primarily aimed at professional film post-production teams, AI short film creators, advertising production companies, and embodied intelligence research teams. For the first two groups, the multi-shot generation and precise editing capabilities can directly shorten the pre-production shooting and post-production special effects workflow periods; for the latter, the Physical AI Checkpoint provides a foundation for world models, assisting robotic systems in perceiving physical motion patterns. In the future, as the community ecosystem matures (with more workflow templates, fine-tuning tutorials, and third-party plugins), LTX-2.5 has the potential to become a benchmark project in the open-source video generation field, encouraging more film teams to incorporate AI into standardized production workflows. Of course, its hardware requirements and 20-second duration limit remain practical constraints that need attention, and we look forward to future versions improving inference efficiency and long-form video generation.

7. Application Scenarios

  • Professional Film Post-Production: In post-production for cinema films, television series, and high-end advertisements, LTX-2.5's native 4K HDR, RAW, and EXR lossless output can be directly integrated into color grading and visual effects software such as DaVinci Resolve and After Effects, eliminating the need for compression and transcoding. The precise video editing capabilities allow post-production teams to rewrite styles, repaint scenes, or replace characters based on real-shot footage, significantly reducing the costs of reshoots and CGI, while maintaining the physical realism of the visuals.

  • AI Short Films and Web Series Creation: The Multishot multi-shot generation capability enables creators to output coherent narrative segments with multiple camera angles and shot compositions in one go, maintaining a high level of consistency in character appearance, environmental atmosphere, and lighting conditions across shot transitions. Combined with the precise video editing feature, teams can adopt a workflow of "real-person action skeleton + AI character replacement and stylized overlay," fundamentally addressing the two major challenges in AI short films: shot continuity and action controllability, significantly improving production efficiency.

  • Advertising and Brand Content: Advertising teams can leverage the Prompt Enhancer to quickly convert product features into cinematic-level prompts, reducing the barrier to prompt engineering. During the creative phase, the Distilled model can be used for rapid iteration and preview, and once the direction is confirmed, the full Dev model can be used to output 4K HDR final videos. The IC-LoRA control mode can lock in brand visual elements (such as specific compositions and product poses), ensuring that the generated content complies with brand guidelines.

  • Physical AI and Robotics Research: The Physical AI Checkpoint provided by LTX-2.5 can serve as the foundational world model for embodied intelligence teams. Researchers use this model to generate video data that adheres to physical laws, helping robotic systems perceive object motion, collision responses, and spatial relationships. This goes beyond pure content creation, offering a source of synthetic data for robot training.

8. FAQ

Q: What hardware configuration is required to run LTX-2.5 smoothly?
A: It is recommended to use a GPU with at least 24GB of VRAM, such as the NVIDIA RTX 4090 (24GB), A100 (40GB/80GB), or higher. With 24GB VRAM, you can generate a 1080p resolution video of approximately 10 seconds. For 4K resolution or longer videos, it is advised to use a GPU with more than 40GB VRAM. For the CPU, it is recommended to use at least an 8-core processor and 32GB of RAM. Currently, AMD GPUs or Apple Silicon with Metal acceleration are not supported.

Q: Does the model support Chinese prompts? How effective is it?
A: The text encoder used in LTX-2.5 is based on Gemma 4 12B, which is primarily optimized for English. While Chinese prompts can be input, the model's understanding of Chinese semantics may not be as precise as with English, especially when it comes to complex action descriptions and camera instructions. For the best results, it is recommended to use English prompts, or use the Prompt Enhancer to convert Chinese shorthand into English cinematic-level instructions.

Q: How is native multi-shot video generation achieved? Is a special prompt format required?
A: Describe the content of each shot in sequence within the prompt, using coherent prose rather than a list of tags. For example: "[Medium shot] A soldier is running in the dusk, with the camera following his back; [Close-up] Switch to his face, with sweat sliding down and a determined look in his eyes; [Wide shot] An explosion rises in the background, with loud sounds." The model will automatically generate multiple shots and maintain consistency in character, environment, and audio during transitions. No additional nodes are required, as LTXVideo workflow already includes native support.

Q: What are the file size and storage recommendations when outputting 4K HDR or RAW formats?
A: A 4K HDR video consumes approximately 500MB–1GB per second (depending on content complexity), and RAW format files are even larger. It is recommended to use a high-speed SSD for storage and reserve at least 100GB of available space. The output path should be set to a non-system drive to avoid disk I/O bottlenecks. During color grading, it is advised to directly import the EXR sequence into DaVinci Resolve or After Effects to preserve full color depth and dynamic range.

Q: Is it possible to preview or interrupt the generation process in real time?
A: In ComfyUI, the generation process is frame-by-frame diffusion, and you can interrupt the task by clicking the "Cancel" button next to "Queue Prompt." However, the model does not support frame-by-frame preview; you must wait for the full generation to view the results. It is recommended to first test the prompt effect with low resolution and short duration (e.g., 512×288, 9 frames), and then increase the parameters for the final generation after confirming the results.

9. Project Links

Related AI Model Articles

© All Rights Reserved. Some content on this site is partially generated by AI with human review.