Adjusted this workflow for a clean single-pass run. Do a second pass only if you like wasting compute.
Key Features
Supports up to 5 reference images.
Audio references for up to 2 speakers.
Preserves multiple characters, clothing, objects, and backgrounds natively.
Slot embeddings and native self-attention retrieval.
Multi-subject and subject-object composition tailored for LTX-2.5.
Prompt Rules (Don't mess this up)
Describe each reference image clearly in the prompt.
Use consistent labels (
Image 1,Image 2, etc.).Define actions and spatial relationships explicitly.
Specify which reference feeds the character, object, clothing, or background.
Match audio references to the correct subject image.
Links & Credits
Character Layouts Workflow: Civitai
MSR LoRA: Hugging Face (Credits to LiconStudio for the solid work on the LoRA).




Base models belong to their respective creators; this is a custom workflow build.
Compatible models support NSFW generation.
Message directly if you run into configuration issues.
If this workflow saves you time: Buy me a coffee