CivArchive
    JoyAI-Echo Multishot Workflow - one character, many shots, same face + voice - v2.0
    NSFW

    Everything here is free and stays free โ€” the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:

    ๐Ÿšจ v2.0 IS HERE โ€” RIFTCAST: CHARACTERS ARE FILES NOW. Design a human from dropdowns, watch them audition, and get a portable character file anyone can reuse โ€” zero training. Plus the 24fps accent discovery that fixes every "why is she suddenly British" bug. Full story in the v2.0 version notes. ๐Ÿšจ


    JoyAI-Echo Multishot โ€” the character video studio for LTX-2.3

    One character. Any number of shots. Same face, same voice, every time โ€” and as of 2.0, your characters are files you can share.

    This is a complete local pipeline for character-driven video on LTX-2.3's joint audio-video model: it generates the picture AND the voice in one diffusion pass โ€” no TTS chain, no lip-sync post, no per-word API costs. A cross-shot memory bank keeps identity and voice locked across an entire multi-shot production, and a set of hard-won fixes makes the base model behave in ways stock workflows can't.

    RiftCast โ€” characters are files now

    Things come through the Rift. Now characters do too.

    A .riftcast cartridge is one file carrying a character's voice (a 4-second anchor clip), face (reference stills), canonical description, and optionally their LoRAs and home environments. Drop it in input/riftcast/, restart, and mention the character in any script: they render with their face and their voice, in any scene, zero training. Cartridges extend the roleplay world's Character Card V3/CHARX lineage into photoreal video โ€” a cartridge can carry a chat persona too, so the same character works in your roleplay client.

    Three ways to get one:

    • Download โ€” WREN.riftcast ships in this package; more on the HuggingFace repo.

    • Cut one from any render you like โ€” one command: python riftcast.py cut <master.mp4> <NAME> <dna.txt>.

    • Design one from scratch โ€” the bundled RiftCast Studio workflow is a full character creator. A Character Designer node covers identity (gender, age, ethnicity, skin, height, build, voice timbre, accent) and a Style + Wardrobe node covers appearance: 71 styles across 10 families, plus hair colour, hair shape, makeup, accessories, demeanor and a freeform wardrobe override. Queue it and the character records an audition tape; the Packer cuts their anchor and reference stills from that render and installs the cartridge automatically. Don't like who showed up? Re-queue with a new seed. Like them? They're a file, forever. (Our test reviewer rated a Designer character's lip sync "Real โ€” very high confidence.")

    One workflow, one switch: RiftCast Studio routes between your prompt files (LPFF/JSON batch rendering, the classic path) and the Character Designer with a single dropdown.

    The style dropdown never sends its own name

    This is the part that makes it work rather than being a word list. "Goth" and "preppy" mean nothing to a video model, so no style label is ever written into a prompt. Each of the 71 entries expands into concrete renderable descriptors โ€” garments named with material and condition, hair shape, makeup with placement, worn objects, and a demeanor that drives how the character physically carries themselves on camera. Picking one goth entry produces, in part:

    "...hair backcombed high at the crown with a straight fringe cut level with the eyebrows, wearing a long-sleeved black velvet dress with a frayed hem over laddered fishnet tights, and buckled boots scuffed grey at the toe, matte pale foundation, black liner drawn thick and winged past the outer corner..."

    Hair colour stays its own dropdown and styles specify only shape, so the two can never collide. Every style carries both a masculine and a feminine wardrobe reading, so it works across presentations instead of being gender-locked, and anything you set explicitly overrides the style's contribution.

    The 24 fps law โ€” why your accents broke

    The single most important thing this package knows: LTX-2.3's joint audio-video prior is 24 fps-native, and render fps is a hidden accent dial. At 25 fps the same prompt and seed render non-rhotic southern British; at 30 fps, broad Australian โ€” and off-24 fps overrides accent wording in your prompt entirely. Verified by A/B with blind phonetic review. Everything here defaults to 24 and warns when you stray. If you want a British or Australian character, render their scenes at 25/30 โ€” it beats any wording. (Pairs with the American-accent audio LoRA, which makes accent wording enforceable in the young-voice registers the base model ignores.)

    The rest of the studio

    • Cross-shot memory bank โ€” identity and voice persist across shots; anchor+latest policy stops drift snowballs.

    • Voice casting โ€” drop a clip in joyecho_voices/<tag>/ and that character speaks with that voice in every render. Script-pinned voices via voice_refs.

    • Finishing โ€” AutoFinish builds your master automatically; deterministic upscale; optional temporal_upscale doubles motion to ~48 fps masters (24 fps render law preserved, audio untouched).

    • Long takes โ€” up to 1441 frames (60 s) single-shot; the old ~10 s lip-sync cliff is fixed at the RoPE-clock level (Bug fix #0).

    • Correctness โ€” a full sampler-path audit (seeded hires, cache keyed on checkpoint, cloned banks), fp8/INT8/GGUF loading paths, and widget values that survive updates (saved by name, not position).

    Hardware

    Built and tested on RTX 5090/3090. GGUF DiT + the 9 GB VAE companion runs the whole stack in ~11 GB system RAM instead of ~60.

    Everything is also on GitHub and HuggingFace โ€” node pack, format spec, LoRAs (surface realism too), quants, and demo cartridges. LTX-2 Community License.

    Description

    ๐Ÿšจ v2.0 โ€” THE BIGGEST UPDATE THIS PACK HAS EVER SHIPPED ๐Ÿšจ

    Characters are FILES now. A character creation screen for real humans. And the discovery of why your voices kept turning British. This isn't a patch โ€” it's a new chapter, and everything in it is free.

    โšก RIFTCAST โ€” YOUR CHARACTER IS NOW A FILE

    Things come through the Rift. Now characters do too.

    One .riftcast file = one complete character: voice, face, canonical description โ€” optionally their LoRAs, home environments, and even a chat persona (CharacterCard V3 compatible!). Drop it in input/riftcast/, restart, mention them in any script, and they show up. Same face. Same voice. Any scene. ZERO training.

    ๐ŸŽ WREN.riftcast ships in this zip โ€” the first video-character cartridge ever cut. Load her and render her anywhere: she arrives with her face, her casual American voice, and her own opinions about laundromats at 2 a.m.

    ๐ŸŽฌ RIFTCAST STUDIO โ€” DESIGN A HUMAN FROM DROPDOWNS

    The bundled Studio workflow is a full character creator now: a Character Designer node for identity (gender, age, ethnicity, skin, height, build, voice, accent) and a Style + Wardrobe node for appearance โ€” 71 styles across 10 families (DARK, ACADEMIC, TECH, SPORT, CLASSIC, STREET, VINTAGE, CUTE, GENRE, WORK), plus hair colour, hair shape, makeup, accessories, demeanor and a wardrobe override. Hit Queue โ€” and your new character records an audition tape. The Packer cuts their voice anchor and reference stills from that very render and installs the cartridge automatically. Designed, auditioned, packed, and castable in about two minutes.

    Don't like who showed up? New seed, new person. Love them? They're a file, forever. (Our blind test reviewer rated a Designer character's lip sync "Real โ€” very high confidence." It thought she was a person. ๐Ÿ’€)

    One switch flips the same graph between the Designer and your classic LPFF/JSON batch prompt files โ€” plus a new Render Clock node: set fps + duration ONCE and every frames/fps socket in the graph follows, with frame counts auto-snapped to legal 8n+1.

    ๐ŸŽญ 71 STYLES THAT ACTUALLY RENDER

    The style dropdown never sends its own name to the model. "Goth" and "preppy" are meaningless to a video diffusion model โ€” so instead every one of the 71 entries expands into concrete renderable descriptors: garments with material and wear, hair shape, makeup with placement, worn objects, and a demeanor that changes how the character physically holds themselves on camera.

    Pick a goth entry and you get "a long-sleeved black velvet dress with a frayed hem over laddered fishnet tights, and buckled boots scuffed grey at the toe, matte pale foundation, black liner drawn thick and winged past the outer corner" โ€” with the word "goth" appearing nowhere.

    Every style ships a masculine AND a feminine wardrobe reading, so nothing is gender-locked. Hair colour is a separate dropdown from hair shape so they cannot collide. Anything you set explicitly beats the style. And the attribute lists grew across the board: age 12, skin 15, ethnicity 30, hair colour 28, hair shape 39, build 13, voice 18, plus new makeup, accessories and demeanor lists.

    ๐Ÿ”ฅ THE 24 FPS LAW โ€” WHY YOUR ACCENTS BROKE (READ THIS!!)

    This might be the most important thing anyone has published about LTX-2.3 audio. The joint audio-video prior is 24fps-native, and your render fps is secretly an ACCENT DIAL:

    • 24 fps โ†’ rhotic General American โœ…

    • 25 fps โ†’ non-rhotic southern British ๐Ÿ‡ฌ๐Ÿ‡ง

    • 30 fps โ†’ broad Australian (rising terminals and all) ๐Ÿ‡ฆ๐Ÿ‡บ

    Same prompt. Same seed. Only the number changes. And off-24, the model ignores accent wording in your prompt entirely โ€” no LoRA, no phrasing, nothing beats the geometry. Verified by A/B with blind phonetic review. v2.0 defaults everything to 24 and warns you when you stray. Flip side: want an authentic Brit or Aussie? Render THEIR scenes at 25/30 and it's free. ๐Ÿคฏ

    โœจ AND THE REST OF THE HAUL

    • ๐Ÿงท Your widget settings FINALLY survive updates โ€” values now save by NAME, not position. Node layout changes can never silently reset your Generate settings again.

    • ๐ŸŽž๏ธ temporal_upscale 2x โ€” LTX's temporal latent upsampler doubles motion to ~48fps masters. Audio untouched, 24fps law preserved. Silk for motion content.

    • โฑ๏ธ 60-second single takes โ€” num_frames cap raised to 1441. The old 481 cap died with the rope-clock bug.

    • ๐Ÿ”ง fp8-mixed gemma encoders load again; the end-of-render console-window flashing is gone forever.

    ๐Ÿ“ฆ UPGRADING

    Unzip over your existing ComfyUI_JoyAI_Echo_GGUF_Nodes, restart, hard-refresh the browser. Open each saved workflow once, set video_fps to 24 (old saves carry 25 โ€” that's the whole accent bug!), confirm your settings, save. Done.

    โ˜• All of this is free and stays free. If it saved you a night of debugging (it contains several hundred of mine): Ko-fi ยท GitHub Sponsors ยท Liberapay ยท or the Buzz tip button right here. โšก

    FAQ

    Workflows
    LTXV 2.3

    Details

    Downloads
    201
    Platform
    CivitAI
    Platform Status
    Available
    Created
    7/30/2026
    Updated
    8/12/2026
    Deleted
    -

    Files

    joyaiEchoMultishotWorkflowOne_v20.zip

    Mirrors

    joyaiEchoMultishotWorkflowOne_v20.zip

    Mirrors