CivArchive
    JoyAI-Echo Multishot Workflow - one character, many shots, same face + voice - v1.6
    NSFW

    Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:

    🚨 v2.0 IS HERE — RIFTCAST: CHARACTERS ARE FILES NOW. Design a human from dropdowns, watch them audition, and get a portable character file anyone can reuse — zero training. Plus the 24fps accent discovery that fixes every "why is she suddenly British" bug. Full story in the v2.0 version notes. 🚨


    JoyAI-Echo Multishot — the character video studio for LTX-2.3

    One character. Any number of shots. Same face, same voice, every time — and as of 2.0, your characters are files you can share.

    This is a complete local pipeline for character-driven video on LTX-2.3's joint audio-video model: it generates the picture AND the voice in one diffusion pass — no TTS chain, no lip-sync post, no per-word API costs. A cross-shot memory bank keeps identity and voice locked across an entire multi-shot production, and a set of hard-won fixes makes the base model behave in ways stock workflows can't.

    RiftCast — characters are files now

    Things come through the Rift. Now characters do too.

    A .riftcast cartridge is one file carrying a character's voice (a 4-second anchor clip), face (reference stills), canonical description, and optionally their LoRAs and home environments. Drop it in input/riftcast/, restart, and mention the character in any script: they render with their face and their voice, in any scene, zero training. Cartridges extend the roleplay world's Character Card V3/CHARX lineage into photoreal video — a cartridge can carry a chat persona too, so the same character works in your roleplay client.

    Three ways to get one:

    • DownloadWREN.riftcast ships in this package; more on the HuggingFace repo.

    • Cut one from any render you like — one command: python riftcast.py cut <master.mp4> <NAME> <dna.txt>.

    • Design one from scratch — the bundled RiftCast Studio workflow is a full character creator. A Character Designer node covers identity (gender, age, ethnicity, skin, height, build, voice timbre, accent) and a Style + Wardrobe node covers appearance: 71 styles across 10 families, plus hair colour, hair shape, makeup, accessories, demeanor and a freeform wardrobe override. Queue it and the character records an audition tape; the Packer cuts their anchor and reference stills from that render and installs the cartridge automatically. Don't like who showed up? Re-queue with a new seed. Like them? They're a file, forever. (Our test reviewer rated a Designer character's lip sync "Real — very high confidence.")

    One workflow, one switch: RiftCast Studio routes between your prompt files (LPFF/JSON batch rendering, the classic path) and the Character Designer with a single dropdown.

    The style dropdown never sends its own name

    This is the part that makes it work rather than being a word list. "Goth" and "preppy" mean nothing to a video model, so no style label is ever written into a prompt. Each of the 71 entries expands into concrete renderable descriptors — garments named with material and condition, hair shape, makeup with placement, worn objects, and a demeanor that drives how the character physically carries themselves on camera. Picking one goth entry produces, in part:

    "...hair backcombed high at the crown with a straight fringe cut level with the eyebrows, wearing a long-sleeved black velvet dress with a frayed hem over laddered fishnet tights, and buckled boots scuffed grey at the toe, matte pale foundation, black liner drawn thick and winged past the outer corner..."

    Hair colour stays its own dropdown and styles specify only shape, so the two can never collide. Every style carries both a masculine and a feminine wardrobe reading, so it works across presentations instead of being gender-locked, and anything you set explicitly overrides the style's contribution.

    The 24 fps law — why your accents broke

    The single most important thing this package knows: LTX-2.3's joint audio-video prior is 24 fps-native, and render fps is a hidden accent dial. At 25 fps the same prompt and seed render non-rhotic southern British; at 30 fps, broad Australian — and off-24 fps overrides accent wording in your prompt entirely. Verified by A/B with blind phonetic review. Everything here defaults to 24 and warns when you stray. If you want a British or Australian character, render their scenes at 25/30 — it beats any wording. (Pairs with the American-accent audio LoRA, which makes accent wording enforceable in the young-voice registers the base model ignores.)

    The rest of the studio

    • Cross-shot memory bank — identity and voice persist across shots; anchor+latest policy stops drift snowballs.

    • Voice casting — drop a clip in joyecho_voices/<tag>/ and that character speaks with that voice in every render. Script-pinned voices via voice_refs.

    • Finishing — AutoFinish builds your master automatically; deterministic upscale; optional temporal_upscale doubles motion to ~48 fps masters (24 fps render law preserved, audio untouched).

    • Long takes — up to 1441 frames (60 s) single-shot; the old ~10 s lip-sync cliff is fixed at the RoPE-clock level (Bug fix #0).

    • Correctness — a full sampler-path audit (seeded hires, cache keyed on checkpoint, cloned banks), fp8/INT8/GGUF loading paths, and widget values that survive updates (saved by name, not position).

    Hardware

    Built and tested on RTX 5090/3090. GGUF DiT + the 9 GB VAE companion runs the whole stack in ~11 GB system RAM instead of ~60.

    Everything is also on GitHub and HuggingFace — node pack, format spec, LoRAs (surface realism too), quants, and demo cartridges. LTX-2 Community License.

    Description

    # v1.6 — Voice casting: your characters' voices are files now
    
    The memory bank keeps a voice *consistent* — but shot 1 always rolled the voice from text, and whatever it rolled, the bank faithfully kept. v1.6 makes the voice a **casting decision**:
    
    **Folder casting.** Drop a clip of your character speaking (4+ seconds, mp4 or wav) into `ComfyUI/input/joyecho_voices/<speaker-tag>/`. Every script whose speaker tag matches that folder gets the voice seeded into the memory bank *before shot 1* — the first shot **continues** your cast voice instead of auditioning a new one, and the same file gives you the same voice in every future render. Replace the file to recast. No widgets, no per-run setup.
    
    **Script-carried casting.** Add `"voice_refs": {"Alice": "path/to/clip.mp4"}` to your script JSON to pin a specific take per character (overrides the folder).
    
    **Audio-only anchors.** A bare wav works — it pairs with the character's ref image from `joyecho_refs/<tag>/` automatically.
    
    **Speaker order derives from your script.** An explicit `"speakers": [...]` array, or just writing `Alice is talking, saying, "..."` in each shot. It travels inside the conditioning (and its cache), so it can't go stale or leak between graphs. The `speaker_order` widget is now only a manual override.
    
    **Anchor + latest.** A speaker's audio context is their anchor plus their most recent shot — one drifted shot can no longer snowball into owning the rest of the video.
    
    **Two-character tip that saves a night:** a2v cross-attention has no spatial addressing — audio at time *t* drives *every* face in frame, no matter how small or distant. Stage **one face per shot** (shot-reverse-shot) and describe only the visible character in that shot's prompt.
    
    ## Fixes (from a full sampler-path audit)
    
    - **Hires refine now respects your seed.** Its re-noise was seeded from a hardcoded constant — every render's refine detail layer was identical regardless of seed, since the feature shipped.
    - **Chained SingleShot graphs no longer condition on the previous queue run's output** (the memory bank was mutated in place through ComfyUI's node cache; incoming banks are now cloned).
    - **Conditioning disk cache is keyed on the checkpoint** — swapping models can't serve you the old model's conditioning. Existing cache entries rebuild once on first render.
    - Cold-start fix: a speaker's first line no longer generates against a silenced bank when per-character memory is active.
    
    ## Upgrading
    
    Unzip over your existing `ComfyUI_JoyAI_Echo_GGUF_Nodes` folder (or replace `nodes.py`, `joyecho_script_picker.py`, `joyecho_prompt_source.py`, and `libs/ltx_distillation/inference/memory_multishot.py`), restart ComfyUI. Existing workflows run unchanged — every new feature activates from your script files or folders, not from graph edits.
    

    FAQ

    Workflows
    LTXV 2.3

    Details

    Downloads
    24
    Platform
    CivitAI
    Platform Status
    Available
    Created
    7/29/2026
    Updated
    8/12/2026
    Deleted
    -

    Files

    joyaiEchoMultishotWorkflowOne_v16.zip

    Mirrors