LTX-2.5 suits teams that want to self-host, fine-tune, or iterate fast rather than call a closed API: Lightricks reports a clip rendered in 6.8 seconds on two Nvidia GB200 superchips. It fits ComfyUI creators and indie studios better than enterprises chasing Sora-level photorealism.
LTX-2.5 is Lightricks' open-weights video and audio generation model, running a 22-billion-parameter dual-stream diffusion transformer with two asymmetric processing streams. It renders synchronized 4K HDR video and audio in one pass, with native multishot scenes and local fine-tuning rights closed competitors don't grant.
Provider: Lightricks · Family: LTX-2
More about Lightricks on HokAI
Input modalities: text, image, video · Output: video, audio
About LTX-2.5
Lightricks released LTX-2.5 on August 11, 2026, out of its LTX division in Jerusalem, Israel, an open-weights checkpoint that generates video and audio together in a single pass rather than compositing them separately. It runs a 22 billion parameter diffusion transformer with two asymmetric processing streams and variable-rate tokenization, and it sits in the LTX-2 lineage, after LTX-2 (October 2025) and the interim LTX-2.3 (March 2026). The release centers on a new rendering approach that spends compute where a scene actually needs it rather than evenly across every frame, paired with an upgraded video decoder and a fine-tuned text encoder built on Gemma 4 12B for tracking multiple subjects, actions, lighting and camera moves through a complex prompt. A lightweight enhancer expands short prompts into fuller cinematic direction automatically, and clip length is inferred from the described action instead of being set as a fixed parameter. The Fast checkpoint covers a wide span, from 720p up through 4K, in landscape or portrait, though anything past ten seconds drops to a lower resolution ceiling; the Pro checkpoint trades that range for a shorter cap and steadier fidelity throughout. A single generation can also produce several connected shots that stay visually consistent scene to scene, and finished output carries wide color and HDR data through the full pipeline for professional grading. Lightricks' own benchmark, run on two Nvidia GB200 superchips, reports a clip finished in 6.8 seconds; the public API takes roughly 23.7 seconds for the equivalent job. Running the full-precision checkpoint locally needs 32GB or more of VRAM, though an FP8 build cuts that by about 40 percent and community GGUF versions target 12 to 16GB consumer cards. Lightricks charges the LTX API by output duration rather than by token, and offers a free self-hosting path for smaller teams under its Community License; larger organizations, or anyone distributing a fine-tuned checkpoint commercially, arrange a separate paid license instead. Weights are distributed through the Lightricks organization on Hugging Face, gated behind license acceptance, with official day-one workflow templates in ComfyUI. A separate Pre-Trained checkpoint is also published as the base model before fine-tuning, for researchers who want to run their own post-training. On Artificial Analysis's text-to-video arena, LTX-2.3, this model's direct predecessor, scored an Elo of 978 (Fast) and 959 (Pro) on the leaderboard that includes audio, well behind category leaders such as Google's Gemini Omni Flash at 1324. LTX-2.5 itself had not been independently scored on the arena in its first week. Lightricks positions the line on speed, open weights and price rather than raw photorealism: Sora 2 and Veo 3.1 lead on physical accuracy and cinematic realism, and Kling 3.0 leads on long, multi-shot consistency. The model card states the checkpoint is not intended or able to provide factual information, may amplify societal biases present in its training data, and may generate content that fails to match a prompt or is inappropriate. Content restrictions run through the Community License's Acceptable Use Policy, which prohibits sexual exploitation of minors, non-consensual sexual content and related categories; the Hugging Face repository gates weight downloads behind license acceptance. Lightricks has not published a named red-teaming partner list or a frontier-lab-style safety framework for the LTX line. Lightricks' LTX Trust Center lists SOC 2 Type II, ISO/IEC 27001:2022 and GDPR alignment at the company level; no LTX-2.5-specific data processing or retention terms beyond the Community License's Acceptable Use Policy were published as of this review.
Pricing
LTX API pricing runs per second of video rather than per token: ltx-2-5-pro ranges from $0.09/s at 720p to $0.37/s at 4K, and ltx-2-5-fast ranges from $0.09/s to $0.30/s across the same span, with no separate request or per-asset fees. Self-hosting the open weights is free for commercial use under $10M in annual revenue under the LTX-2.x Community License.
Key Features
- Native Multishot: A single generation can produce multiple connected shots that hold character, environment, lighting and voice consistent across cuts.
- Diffusion Fidelity Rendering: Allocates rendering compute by scene complexity instead of spending it evenly, the core architectural change behind this release's speed gains.
- 4K HDR Output: Renders up to 4K at 50fps with ACES EXR in/out and 16-bit linear range preserved through the pipeline, built for color-finishing workflows.
- Auto Duration: Predicts the appropriate clip length from the described action before diffusion starts, instead of requiring a fixed duration input.
- Synchronized Audio-Video Generation: Produces matching soundtrack and dialogue alongside the picture in one diffusion pass, instead of layering on a separate audio model afterward.
- Open Weights With Fine-Tuning Rights: Ships as an open checkpoint on Hugging Face with day-one ComfyUI support, free to fine-tune locally under the Community License.
Pros
- Free for smaller studios to use commercially under Lightricks' Community License, plus local fine-tuning rights closed video APIs don't grant.
- Native multishot generation holds character, lighting and voice consistent across cuts in one pass, a feature usually reserved for closed cinematic tools.
- Day-one ComfyUI and Hugging Face support with official workflow templates, lowering the setup barrier for local generation.
Cons
- Predecessor LTX-2.3 trailed the video arena leader by a wide Elo margin on Artificial Analysis, and LTX-2.5 had no independent arena score in its first week.
- Full-precision local inference needs 32GB+ VRAM; most consumer GPUs require a quantized GGUF or FP8 build with some quality trade-off.
- Organizations past the self-hosting revenue cap, or anyone distributing fine-tuned checkpoints commercially, need a separate paid license from Lightricks.
Benchmarks
Frequently Asked Questions
How much does LTX-2.5 cost to generate video?
Running LTX-2.5 through the LTX API is billed by the second, not the token: the Pro checkpoint starts at $0.09 for a second of budget-resolution output and rises to $0.37 at full 4K, while the Fast checkpoint tops out lower, at $0.30. Teams that would rather skip per-second charges can self-host the open weights for free under Lightricks' Community License, provided their annual revenue stays under the free-tier threshold.
How does LTX-2.5 compare on benchmarks vs Kling 3.0?
LTX-2.5 had not been independently scored on Artificial Analysis's video arena in its first week, while Kling 3.0 posted an Elo of 1,248 on the same leaderboard. LTX's prior release scored lower still on that board. LTX competes on generation speed, open weights and price rather than arena ranking.
Is LTX-2.5 open source or proprietary?
LTX-2.5 is open-weights rather than fully open-source: Hugging Face hosts the checkpoint behind a license click-through, and Lightricks' Community License lets smaller organizations run it commercially at no cost. Larger companies, or anyone reselling a fine-tuned version, arrange a separate license directly with Lightricks.
Does LTX-2.5 train on user data?
Lightricks has not published a specific training-on-inputs policy for LTX-2.5's API outputs. A related LTX Platform privacy policy, last updated May 10, 2026, covers a different product (Face and Voice Models) and states that data is deleted within two years of last use; broader retention terms for LTX-2.5 generations are not separately disclosed.
Who is LTX-2.5 best for and who should avoid it?
LTX-2.5 fits indie filmmakers, ComfyUI workflow builders and small studios under the $10M revenue threshold who want fast, fine-tunable, locally-run video generation. Enterprises chasing guaranteed photorealism should look at Sora 2 or Veo 3.1 instead, and studios needing long-form multi-shot consistency should consider Kling 3.0.