Generate video with matching sound from text or a photo
Made by LTX.io
View profile on Hugging Face (opens in a new tab)LTX-2.5 is a 22-billion-parameter open-weights model from Lightricks that generates video and sound together: describe a scene in words, or hand it a photo, and it produces a clip with a soundtrack in one pass, so picture and audio stay in sync.
New in this version are multi-shot clips — two to four connected shots in one generation — and upscalers that push the result toward 4K. It runs on your own hardware through ComfyUI or Python, but the files are tens of gigabytes and need a powerful graphics card. The weights require a free Hugging Face account and accepting Lightricks’ community license, which restricts commercial use for larger companies.
Can I use this?
You'll need A powerful GPU, ComfyUI or Python, and tens of gigabytes of disk · Setup needed
Worth knowing The weights are gated: you need a free Hugging Face account and must accept Lightricks’ community license before downloading, and that license restricts commercial use for larger companies — read it before business use. Speed and quality claims are Lightricks’ and the community’s own. BuildTube has not run or verified this model.
Why it's here
- 5,805 people have liked it on Hugging Face.
- It was downloaded 1,588,619 times in the last 30 days.
Numbers from the snapshot taken 1 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?