View complete timeline
October 23, 2025Model Release
LTX-2Lightricks

Nine days after Veo 3.1, Lightricks — the Jerusalem company behind Facetune and Videoleap — shipped the first open model that generates picture and sound together in a single pass. Not video with audio dubbed on afterward: an asymmetric dual-stream diffusion transformer, 19B parameters split 14B video and 5B audio, producing dialogue, lip sync, and ambience jointly with the frames. Native 4K at up to 50 fps, roughly 10-second clips, and it runs on a consumer GPU with 12 GB of VRAM. Native audio had been the closed models' moat since Veo 3 in May; LTX-2 was the first time you could have it on your own hardware. The honest asterisk is that on launch day you couldn't. Lightricks announced LTX-2 as open source with weights promised "later this fall," and shipped only inference tooling and a paid API — $0.04, $0.08, and $0.16 per second for Fast, Pro, and Ultra. The weights actually landed on January 6, 2026, alongside the full training code, NVFP8 quantization, and ComfyUI support, under the LTX-2 Community License: free for academic work and for companies under $10M ARR. Zeev Farbman, co-founder and CEO, framed the launch with a line the model mostly earned — "diffusion models are reaching a point where they no longer just simulate production, they are production" — though early users found the usual seams in lip sync, crowded shots, and any frame containing text. It entered the Artificial Analysis image-to-video board at #3 and was the highest-ranked open model by early 2026, becoming the base that LTX-2.3 and LTX-2.5 were built on.

Native 4K · ~10s at up to 50 fps · 19B (14B video + 5B audio)