✨ Unlimited length, full lip-sync, and 4K scaling on a budget? Yes, it’s finally here. 🤝
Forget short, silent clips. This advanced LTX 2.5 workflow breaks the barrier by combining recursive looping pipelines with high-fidelity audio syncing. Whether you are generating full-length music videos or deep dialogue scenes, this pipeline handles extended runtimes smoothly without crushing your hardware.
The best part? It integrates INT8 quantization to keep things lightning-fast, finishing off with a massive RTX upscale for production-ready quality.
🛠️ Key Features in V1.0
Seamless frame-looping logic added to keep transitions flawless. Audio-to-lip matching is locked frame-by-frame. The integrated RTX upscaler now actively preserves micro-details and facial expressions throughout the entire length of the video.
🎨 The Feature Stack
Infinite Video Generation: Uses smart looping technology to create continuous, full-length content instead of short bursts.
Flawless Lip-Syncing: Specifically engineered for heavy dialogue, cinematic acting, and music videos.
Optimized INT8 Pipeline: Massive VRAM savings courtesy of heavily compressed, distilled architecture.
RTX Spatial Upscaling: Standard generations are automatically upscaled into razor-sharp, cinematic resolutions.
Enhanced Visual Clarity: Injected with edge-enhancement LoRAs to eliminate typical AI blur and keep textures crisp.
🧠 The Shopping List: Required Models
To prevent execution errors, make sure you have these exact files placed in your ComfyUI directories:
⚡ The LTX 2.5 Core Engine
Transformer:
LTX2.5-22b-distilled-transformers-comfy-int8-convrot.safetensorsText Encoder:
gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensorsVideo VAE:
LTX2.5-video-vae-conv-bf16.safetensorsAudio VAE:
LTX2.5-audio-vae-bf16.safetensors
🌀 The Enhancer & Upscale Stack
LoRA:
LTX2.3_Crisp_Enhance.safetensorsSpatial Upscaler:
LTX-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors
🛠️ Quick Start
Drag and drop the workflow file directly into your ComfyUI workspace.
Verify your model paths point correctly to your INT8 checkpoints and text encoders.
Load up your audio track, Queue Prompt, and watch full-length synced video come to life!
💬 Join the Cinematic Movement!
If this pipeline helps you create your next masterpiece, smash that Like button and drop a review!
👇 Got a music video or dialogue test? Show off your 4K RTX-upscaled creations in the comments below! 👇
Description
Comments (8)
Cant wait to try this 👍🏼
Epic workflow 🔥
Mind blown 🤯👍🏼
Great WF.
Should it be possible to add an option for written text in a box to be spoken with voice sample ?
Thanks
Voice cloning with custom text ? I'll definitely look into this, it seems possible, perhaps another audio model can be infused to help get this done. Lets see if it works. If it does, V2 is going to have it. You're welcome.
@BopStar thanks. Your actual V1 works perfectly.
It breaks if you set the switch to the mode without audio trim.Hi, thanks for your comment.
There are two ways. Ensure SET_AUDIO remains behind Load Audio node and you can bypass Audio Trim. Other way is even if you don't need to input custom start. Just keep the audio trim on & put start_time as 0.00 (default). Ensure you have loaded the required audio as well. Works perfectly in both situations, I tested it again.