First open-source production-ready AI model for synchronized 4K video and audio generation
LTX-2 is Lightricks' second-generation open-source video foundation model, built on a 19-billion-parameter Diffusion Transformer (DiT) architecture. It is the first open-source model to combine four capabilities simultaneously: 4K resolution output, 50 fps frame rate, audio synchronization, and clips up to 20 seconds — a combination that places it in direct competition with commercial-grade video models on technical specification alone. LTX-2 supports text-to-video and image-to-video workflows, generating both the visual content and a synchronized audio track (ambient sound, narration cues, and environmental audio) in a single pass. This eliminates the common workflow of generating silent video and separately producing or sourcing audio. The model achieves a 3x performance improvement on NVIDIA RTX 50 series GPUs, and Lightricks has published optimizations for the RTX 40 series as well, making 4K generation feasible without data center hardware. LoRA fine-tuning support allows researchers and studios to train custom styles and character appearances into the model. LTX-2 is released under an Apache 2.0-compatible license, with full model weights available on Hugging Face and all training code on GitHub. The dedicated interface at ltx-2.ai and the broader LTX ecosystem at ltx.io provide non-technical users web access without self-hosting. For professional production, LTX Studio at ltx.studio offers a full creative suite built on top of LTX-2.
Explore other tools in this category
All top AI video models in one platform — generate from text, image, or video
Generate 1080p cinematic videos up to 15 seconds from text or images
Google DeepMind's cinematic video model with native synchronized audio
ByteDance's dual-branch AI video model with natively synchronized cinematic audio
Create cinematic AI videos from text and images in a unified browser-based workflow
xAI's free AI video generator — turn images into short videos with Grok