Workflow generates audio from video using LTX-2.3 model.
The workflow is tested on recent ComfyUI version and recent versions of nodes with configuration PyTorch 2.9.0+Cu130, Python 3.13.11, RTX5090, Windows 11.
The workflow is optimized for input videos in 1280x720 resolution and 25 fps.
Description
Simple parameters. Possible future developing.
FAQ
Comments (6)
Where can I find the alliterated text encoder? Is it necessary or will the normal one work?
But be aware, there is mistake in classification - the file to download is not a lora, the file is text enconder indeed.
The normal will work.
CFG=3, Image strength=0.12 gives even better results.
I have a lot of white noise with this, any solutions?
