CivArchive

    Support

    Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:

    Type a story. Get one continuous video, with sound. Multi-shot scenes render as a single take - no last-frame chaining, no quality loss from shot to shot. That's the whole pitch.

    Both workflows now read left to right: numbered lanes, and you only ever touch lanes 2-4. Everything below the main row is optional.

    What you need

    • ComfyUI + this node pack (Manager: MiniMax-H3 Multishot, or the zip on this version).

    • A MiniMax-H3 checkpoint (links on this page). 24 GB card? Take a GGUF.

    SEAMLESS CHAIN - multi-shot scenes as one take

    Lane by lane:

    • README - the quick start lives on the canvas itself.

    • 1 - MODELS - pick your H3 checkpoint and text encoder. VAEs are preset; LoRA slots are empty until you fill one.

    • 2 - ANCHORS (optional) - a photo to open shot 1 on (enable its gate), and a short voice clip to lock the speaker's voice.

    • 3 - YOUR PROMPTS - type your idea in the box, or point the switch at a prompt file. The writer expands it into shot prompts. Writing your own? Set the writer to passthrough (raw JSON, skip LLM) and paste shots separated by --- lines.

    • 4 - CONTROLS - size, frames per shot, steps, and take_seconds (total length; 30 is a good first run). The switches stay off unless you installed the pack a switch names.

    • 5 - ENGINE - nothing to change. The remote encoder lives here if you want the text encoder on a second PC: enter its address, flip the encoder switch, free ~15 GB.

    • 6 - OUTPUT - your video and its audio save here.

    Optional panels below the main row: reference images (your character, ref2va checkpoints - folder per character + AUTO REFS on), V2V reference (a clip whose look guides the render), FFLF plates (flf_chain mode only), audio spine (a soundtrack the take follows).

    EXTEND TAKE - one person talking, as long as you want

    Same lanes, different job: one premise becomes ONE continuous speech cut across windows.

    • 1 - MODELS - same as above.

    • 2 - ANCHORS - a photo of your speaker (shot 1 opens on them) and a voice clip. More useful here than anywhere: one person carries the whole take.

    • 3 - YOUR PROMPT - ONE premise, one speaker. The writer writes the whole speech. num_shots 0 = it decides. Passthrough works here too.

    • 4 - CONTROLS - take_seconds is the star: 30 ships, 60 clears TikTok's minute. window stays on auto - it sizes itself to your card.

    • 5 - ENGINE / 6 - OUTPUT - same as above.

    Keep takes to about 4 windows for now - very long takes slowly sharpen.

    Rules of thumb (both workflows)

    • Spoken lines: 8-12 words per shot. Short lines sync; long lines garble.

    • Say the sounds you want ("rain on the roof, a fridge hum") or it invents its own.

    • Keep your character's face in frame - faces carry identity between shots.

    If something breaks

    • Red node? Update the pack in Manager, restart, reload the workflow from disk.

    • Render crawls at low wattage? Lower resolution or frames per shot, or use the remote encoder.

    • Only one of the two workflows shows in your sidebar? Fixed in 2.6.5 - re-download both.

    • Still stuck: comment with your console log. I answer.

    Deep dives: the two articles linked on this page. Every lane also has a short note on the canvas.


    Detailed guide for people that can read good:

    Every setting explained: the Seamless Chain deep manual | Civitai

    Description

    Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:

    ⚡ Or right here: the Civitai tip button on this page sends Buzz directly

    The joins are gone. Two independent video-understanding models were shown a four-shot chain from this version, blind and with no context. Both described it as one continuous unedited take - no cuts, colour consistent, audio unbroken.

    That is the whole point of 2.0. Everything below is either what made it possible or what got cleaned up on the way.

    A second way to join shots

    Chaining used to mean handing the next shot the previous shot's last frame. That still works, is still shipped, and is still what the CORE workflow uses. 2.0 adds continuity = context_pin, which is a different mechanism entirely.

    The previous shot's last 22 frames ride into the next shot as raw latents - bit-identical, never decoded to pixels and re-encoded - placed at interior keyframe coordinates, with a timeline-placed audio reference alongside them. The regenerated head is trimmed on decode. Colour, motion and voice cross the boundary as data instead of as a description, which is why the join stops being a place where things can drift.

    This mode needs the ComfyUI-H3-Motion-Context pack. Without it, set continuity to first_frame - the model's own trained hand-off, no extra install, and what the CORE workflow ships with.

    Seam audio that stops eating words

    The boundary audio cut used to land blindly at the start of the incoming shot, which shaved the attack off any word the model happened to place there - audible as a tiny blip at a join. It now lands in the quietest gap within the shot's first 0.75 seconds, then welds with the same 40 ms equal-power crossfade. A line that begins early survives intact.

    Voice and identity anchors, on the chaining sampler

    The memory-bank sampler now takes the same anchor inputs the single-shot sampler always had: reference_images, voice_ref, self_anchor_voice, reference_image_size, preview_first_shot, two_pass_upscale and its three dials, plus sampler_override / scheduler_override so one master panel can drive the sampler and scheduler by name.

    The bank already carried voice - each slot is a video reference with the audio under it - but only from shot 2 onward, because shot 1 renders against an empty bank. self_anchor_voice closes that gap: shot 1's own rendered voice becomes <Audio 1> for every later shot. Operator-supplied references are placed ahead of the bank slots, so <Picture n> and <Audio n> numbering stays fixed as the bank fills and a prompt's bindings cannot drift mid-chain.

    Two-pass upscale cannot be combined with continuity = context_pin or latent_handoff, or with an audio spine. Those carry the previous shot's raw latents - or one locked denoise trajectory - across the join, and a two-pass render preserves neither. The node stops with an error naming the conflict rather than quietly producing a weaker join. Two-pass is available on cut, seamless, seamless_tail, first_frame and flf_chain.

    A VRAM and speed panel

    One panel, three lazily gated model patches - memory-efficient attention, chunked feed-forward, and a block cache - plus an activation-reserve control. Every switch off reproduces the shipped recipe exactly, and a gate that is off means the patch node never executes at all, so there is no cost to leaving them alone.

    And a reserve that understands payload

    Memory measurements are now keyed by shape and conditioning payload, so a bare first shot and a reference-laden later shot no longer share one estimate. Measured pools are no longer overridden by a fixed floor, an unseen payload variant estimates from a measured sibling, and a spill into system RAM is now detected and named in the console rather than presenting as an unexplained five-times slowdown with nothing in the log.

    The writer knows the join rules now

    The LLM prompt writer gained a join_style control. Set it to a chained style and the render-verified boundary rules are appended to its system prompt automatically - open each shot holding the previous shot's closing arrangement, land settled, never straddle a line across a boundary, repeat every description verbatim. Generated scripts obey them without you memorising anything. Hand-written scripts still need to follow them; they are documented in PROMPTING.md with a worked four-shot example.

    Three workflows instead of five

    • H3_Seamless_Chain_v2 - everything, with the optional lanes gated off by default: master controls, LLM writer, VRAM panel, identity and voice anchors, an episode/batch prompt source, FFLF boundary plates, and an audio spine.

    • H3_Seamless_Chain_CORE - the same job with zero third-party packs. Type shots into the script box and queue. Start here.

    • H3_Keyframes - unchanged, and still its own thing: single clip, anchors at chosen frame positions, per-anchor condition strength.

    H3_Multishot_AIO and H3_Multishot_MEMORY retire. Every lane they had is in v2 - the AIO's episode source, plate chain and audio spine were folded in, and MEMORY had nothing v2 lacks. Your existing copies keep working; there is just no longer a reason to open them.

    Smaller things that were worth fixing

    • flf_chain selected with no boundary plates wired now stops with a clear error instead of quietly rendering an unanchored chain.

    • Interior keyframe anchoring and the Motion Context pack no longer fight over the same patch site. If that pack is installed it owns the site and this one stands down with a line in the log; if it is not, this one fills the gap. First and last anchors work either way.

    • GGUF vision sidecars can be named explicitly. ComfyUI-GGUF pairs the -mmproj file to a GGUF text encoder by filename, in the encoder's own folder only - rename either, or split them up, and it loads the encoder without its vision tower, which presents as the model ignoring your reference image. This pack's CLIP loader now raises instead of continuing blind, uses the only mmproj beside the encoder when there is exactly one, and takes an mmproj_name widget so you can point at the file directly - so the filename rule documented on the encoder listing stops being something you have to obey.

    • The pack now imports defensively - one module failing no longer takes every node in the pack down with it.

    Shipped defaults

    ref2va checkpoint, continuity = context_pin, voice self-anchor and the identity bank on, euler + beta57, 14 steps, 362 frames per shot (~15.1s), 24 fps, every VRAM switch off.

    Blind review ran on the lighter configuration - fl2va, no voice anchor - which is the same chaining mechanism with fewer reference rows. ref2va ships as the default because it makes voice and character identity explicit rather than emergent; switching to the reviewed path is two changes, documented in SETTINGS.md.

    Honest limits

    • Audio dulls slightly per hop on very long chains - restart the chain on a scene cut, where a fresh start costs nothing.

    • Resolution cannot change mid-chain.

    • flf_chain has not been rendered against a fully colour-matched plate set.

    • Two-pass upscale is off by default and excluded from the raw-latent continuity modes (above), so the seamless chain has no built-in upscale path yet. Upscale after the fact, outside the graph.

    • Keep the mux at 24 fps. Other rates audibly shift voice accents; that is the model's audio lane, not the muxing.

    FAQ

    Comments (39)

    snake88Aug 11, 2026· 1 reaction
    CivitAI

    the AIO script seems to be closer to what I want, the ref2a flexibility is really good and I trust it to preserve what I want at the seams more, the only thing lacking is sometimes tricky to get it to seamless transition instead of jump into slightly different positions.

    joeygambino
    Author
    Aug 11, 2026· 1 reaction

    Try latest v2.1 - seamless transitions built in by default.

    LemmingWolf01Aug 11, 2026· 1 reaction
    CivitAI

    Thank you for all the work you are putting into these workflows, it's appreciated.

    Unfortunately I'm have difficulty with v2.0 as I can't find JoyEcho_LLMEnhance or JoyEcho_PromptSource anywhere. Of course, they just so happen to support the feature I'm eager to try. Any pointers?

    Update1: I've installed JoyEcho_LLMEnhance from RealRebelAI's ComfyUI_JoyAI_Echo_GGUF_Nodes pack. Still looking for JoyEcho_PromptSource

    Update2: I had to drop joyecho_prompt_source.py from HF joeygambino/joyai-echo-multishot-workflow into custom_nodes\Comfyui_custom_scripts folder. I don't know if it's vital but I removed the second underscore in the file name as that was the name thrown by error in Comfyui.

    I think I'm ready to go, I'll leave this here in case it helps anyone else.

    joeygambino
    Author
    Aug 11, 2026

    v2.1 going up shortly to fix some bugs I didn't catch locally and will take care of this. Sorry!

    vladulidloAug 11, 2026
    CivitAI

    Both v2.0 and 2.1 do not seem to work with continuity=context_pin even with ComfyUI-H3-Motion-Context installed.

    [WARNING] h3_motion_context: another pack has already patched MiniMaxH3.extra_conds (it now comes from '/home/vlady/apps/ComfyUI/custom_nodes/ComfyUI-H3-Multishot.h3_avbank_probe'). Both packs are solving the same keyframe/ref collision and they cannot both own it, so this one is refusing. Disable one of them and restart.

    [ERROR] !!! Exception during processing !!! h3_motion_context: the payload patch could not be applied. Without it the audio ref would overwrite the pinned video latents and the motion context would be lost. The reason was logged just above this error.

    vladulidloAug 11, 2026

    A different issues found:
    continuity=seamless behaves as a cut, not seamless at all.

    continuity=seamless_tail errors out with T2V mode AFTER sampling, not before:
    [ERROR] !!! Exception during processing !!! only first/last keyframe anchors are supported

    I'm in search of seamless T2V (and I2V) clip chaining and could not find working setting in the current version. Both identity anchor gate [OFF] and FFLF PLATES gate [OFF - flf_chain only] are set to T2V

    joeygambino
    Author
    Aug 11, 2026· 1 reaction

    @vladulidlo All three confirmed, and thank you - this is an excellent report. 2.1.1 is up with the fixes.

    context_pin + Motion-Context: my pack was grabbing the same patch site Motion-Context needs, before their pack could. Their code publishes a compatibility marker for exactly this situation; mine now honours it, so the two coexist and load order no longer matters. Verified with a live context_pin render.

    seamless_tail: real conflict - it needs interior keyframe anchors, which collide with Motion-Context's ownership of that patch math. It now stops before sampling with a clear message instead of dying after your first shot. With Motion-Context installed, use context_pin - it's the stronger mechanism and what that pack is for.

    seamless: you're right, and the tooltip now says so - it's a legacy latent-only soft pin kept for comparison, and it often reads as a cut. For seamless T2V chaining use context_pin (or first_frame on an fl2va checkpoint). Both are the measured, working paths.

    The bug never showed on my machine because of an install-layout difference that disabled the conflict detection - also fixed, and my release testing now runs on a packaged clean install so this class doesn't slip through again.

    vladulidloAug 11, 2026

    @joeygambino 

    Thank you! For both the fixing and fixing it so quickly!

    Emanresu_ymAug 11, 2026· 1 reaction
    CivitAI

    Seems like the bugs started eating into the bugs, chill, don't rush, take your time, customers can wait.

    joeygambino
    Author
    Aug 11, 2026

    Ha, thanks. I do tend to rush when I have a new feature to show off. A lot of bugs don't pop up until someone reports them, because the workflows are functioning perfectly for me, but then I realize the things other people just don't have installed.

    Emanresu_ymAug 11, 2026· 1 reaction

    @joeygambino yeah people are too excited for new model and what they can do. also samples looks good, compared to previous degrading over time was visible, now it looks stable through all 30 secs. will be trying lastest WF later on, good job!

    vladulidloAug 11, 2026· 1 reaction
    CivitAI

    I would like to point to a fork https://github.com/ethanfel/ComfyUI-MiniMaxH3-Contex-Loop

    ethanfel's fork is 51 comits ahead of original https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context, but also not (yet) compatible with ComfyUI-H3-Multishot.

    So if you are seeing:

    RuntimeError: continuity=context_pin needs the ComfyUI-H3-Motion-Context pack installed (github.com/NikoDemon80/ComfyUI-H3-Motion-Context)

    You could have downloaded the fork instead of https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

    joeygambino
    Author
    Aug 11, 2026· 1 reaction

    Thank you - you're right, and I've verified it against the fork's source. It registers 18 node ids (MiniMaxH3LoopTrim, the MiniMaxH3Chain* family, the Scheduled* reference nodes) and deliberately does not re-register MiniMaxH3MotionContext. From its own init.py: "The original Motion Context, Save Latent, and Load Latent ids remain exclusively owned by Niko's upstream pack."

    So, the fork is a complement rather than a replacement - install both. They're built to coexist, and this pack works with one's runtime patches because all three honor the same patch-ownership markers.

    My error message was unhelpful about that, so 2.1.1 now detects the fork and says exactly this instead of just naming a repo you thought you'd installed. It's also documented in INSTALL.md.

    That fork looks well worth a look on its own merits, by the way, a disk-backed chain/loop system with review gates and checkpoint resume is solving a different problem than this pack and solving it further than I have.

    vladulidloAug 11, 2026

    @joeygambino Good to know! I'm glad you find the forked repo interesting as I do.

    ProfugoBarbatusAug 11, 2026· 3 reactions
    CivitAI

    Just installed 1.5 last night, after wondering to myself "This new H3 stuff is amazing, just wish I could load more than start/stop keyframes" and looking at civitai for a lark. Blown away by this stuff - once 2.0 or its successors settle out from the bugs, I'll update, but this is already incredible, bordering on revolutionary for me to mess around with. Well done!

    joeygambino
    Author
    Aug 11, 2026

    Thank you! 2.1.1 should mostly be bug free now - but if you find anything, I try to be quick about fixing things. Sometimes stuff that works perfectly on my machines, don't necessarily work well on others due to different Comfy versions, hardware, node packs installed, etc. I don't know something is broken until someone tells me.

    Emanresu_ymAug 11, 2026· 2 reactions
    CivitAI

    I have noticed that this larryvrh/MiniMax-H3-Turbo-Lora is not burning the output like the lightning one does, anyone else been using it with this workflow?

    joeygambino
    Author
    Aug 11, 2026

    The only one I've used so far is minimax_h3_turbo_4step_ckpt500.safetensors

    And I can't even say I've tested it enough to know it works. When I did test it, it didn't seem to work very well with my workflows unless it was at 12 steps + only Euler - which at the time it was in Alpha so I didn't bother testing further and figure people will use their own anyway. I didn't realize he'd put a hundred more options up though, so I may have to give them a shot.

    I've tried to optimize things enough so you don't need a lightning/turbo lora, and 4 steps on a video render seems nuts to me, but I suppose I should give it a shot.

    Emanresu_ymAug 11, 2026· 1 reaction

    @joeygambino with latest V4 he even says 4 steps is not enough need to be minimum 6-8, at 8 no improvement would be gained. he also uses custom sampler and custom lora loader for those who has less Vram.

    audrymax919Aug 11, 2026· 2 reactions
    CivitAI

    Thanks for all the work on this pack, the chaining is genuinely great.

    Found a bug though: guide_audio (Audio Spine) with a real voice track outputs static/hiss instead of the audio, on ref2va. Same audio file works perfectly through the native MiniMaxH3ReferenceToVideo node, so it seems isolated to the Audio Spine injection path

    joeygambino
    Author
    Aug 11, 2026

    On it, next update, coming tonight.

    audrymax919Aug 12, 2026

    @joeygambino Thanks for the reply! Quick update: I've now tested up through 2.1.6, same result.

    The resample fix from 2.1.3 is confirmed working on my end (console shows the resample happening correctly), but the audio is still coming out as static/garbled, no change from before the fix.

    I've ruled out continuity mode (context_pin/latent_handoff), checkpoint format (safetensors and GGUF), scheduler (beta/beta57), and LoRA, same result every time. voice_ref works perfectly on the exact same file, so it's specifically the guide_audio path that's still broken for me.

    How are you testing this on your end? Trying to figure out what's different about my setup (RTX 5090, ComfyUI 0.32.0).

    joeygambino
    Author
    Aug 12, 2026

    Got it — and the answer to "how are you testing this on your end" is the bug. I am on ComfyUI 0.30.0. You are on 0.32.0. It works here and cannot work there, and that is entirely on me for not testing across versions.

    What changed. 0.32.0 introduced ModelSamplingAV, and ComfyUI now carries the audio half of the audio+video pack scaled onto the video schedule:

    process_latent_in audio slice x (shift / audio_shift)

    process_latent_out audio slice x 1 / (shift / audio_shift)

    For H3 those shifts are 12 and 3, so the audio latent inside the sampler lives in a 4x-scaled domain. The Audio Spine locks its encoded audio into that pack during sampling — and it was writing raw, unscaled values. Every locked column lands 4x too small, which decodes as exactly the static you are hearing. On 0.30.0 there is no such scaling, so raw was correct.

    That accounts for everything you found, and your process is what made it findable:

    - voice_ref works on the same file — it goes through the text conditioning and never touches the sampler's latent, so the scaling never applies to it.

    - The 2.1.3 resample fix fires and changes nothing — you were right, it works. Encoding was never the problem; the problem is one step later.

    - Continuity mode, checkpoint format, scheduler, LoRA all irrelevant — none of them touch this path. Ruling them out is what pointed at the sampler.

    Also worth knowing: *audio_lock has the identical bug** on 0.32.0, same code path, same cause. It is fixed by the same change.

    The fix reads the scale off the live model_sampling object rather than hardcoding 4, so it stays correct if you change the shifts with MiniMaxH3SigmaShift, and it leaves 0.30.0 behaviour byte-identical. It is written and deployed on my side but I have not render-verified it on 0.32.0 yet - I am setting up a 0.32.0 instance to reproduce your exact failure and confirm the cure rather than ship it on code reading alone. It will be in the next release, which is close.

    Until then, honestly, there is no clean workaround. voice_ref will hold one voice across shots and is the nearest thing, but it is not the spine — it does not lock every shot to one continuous performance. If you need the spine specifically, 0.30.0 is the only place it currently works, and I would not recommend downgrading a whole install for one feature when the fix is coming.

    Thank you for staying with this through 2.1.6 and for testing so carefully. Four ruled-out variables plus "voice_ref works on the same file" is what turned this from a shrug into a one-line fix.

    joeygambino
    Author
    Aug 12, 2026

    Oh, and, sorry about "coming tonight" - I got sidetracked trying to add too many features at once, which I tend to do. I am going to say it again though... fix is coming tonight (I hope).

    audrymax919Aug 12, 2026

    Thanks for the deep dive, really appreciate it! No worries about the delay, I'll wait for the fix and try guide_audio again once it's out

    egin1992654Aug 11, 2026
    CivitAI

    attention (gated) and chunk (gated) nodes dont work for me

    joeygambino
    Author
    Aug 11, 2026

    Can you paste the errors from the terminal?

    egin1992654Aug 12, 2026

    @joeygambino [WARNING] invalid prompt: {'type': 'missing_node_type', 'message': "Node 'attention patch (gated)' has no class_type. The workflow may be corrupted or a custom node is missing.", 'details': "Node ID '#9'", 'extra_info': {'node_id': '9', 'class_type': None, 'node_title': 'attention patch (gated)'}}

    egin1992654Aug 12, 2026

    @joeygambino also in multishot wflow now have error
    Prompt outputs failed validation: H3MultishotMemorySampler: - Value 4 bigger than max of 3: memory_frames

    Это может быть связано со следующим скриптом:
    /extensions/comfyui-easy-use/assets/extensions-WrZZZUnM.js

    joeygambino
    Author
    Aug 14, 2026

    @egin1992654 Sorry for the late response, I missed you replied.

    ## 1. The gated nodes: two packs to install

    The full workflow uses three nodes from two packs I am not allowed to bundle.

    I shipped them switched off, assuming that was enough - it is not.

    ComfyUI checks that every node class exists before it will queue anything, even a node that is switched off, so a missing pack stops the whole workflow instead of just that one feature.

    Install these two and the workflow runs exactly as shipped:

    ComfyUI-sol-attn (provides two of the three)

    https://github.com/Saganaki22/ComfyUI-sol-attn

    comfyui-minimax-h3-blockcache-T8 (provides the third)

    https://github.com/T8mars/comfyui-minimax-h3-blockcache-T8

    Via ComfyUI Manager (easiest): Manager > Install via Git URL, paste each URL in turn, then restart ComfyUI.

    Or by hand:

    cd ComfyUI/custom_nodes

    git clone https://github.com/Saganaki22/ComfyUI-sol-attn

    git clone https://github.com/T8mars/comfyui-minimax-h3-blockcache-T8

    then restart ComfyUI. Check the console on startup - if either pack fails to import it will say so there, and that message is the thing to send me.

    Reload the workflow afterwards. The three nodes will resolve, and the error goes away. They are speed and memory optimisations, so you will also get a faster render out of it.

    One thing worth knowing: those three ship bypassed on the canvas. Installing the packs stops the error. If you then want the speed as well, select each node and press Ctrl+B to un-bypass it, and turn on the matching switch on the FEATURE SWITCHES panel. Leaving them bypassed is fine too - everything renders identically, just slower.

    ## 2. The memory_frames error

    Value 4 bigger than max of 3 - that dial only accepts 0 to 3.

    Open the H3MultishotMemorySampler node and set memory_frames to 0 (that is the shipped default), then queue again.

    If other dials on that node also look wrong, the workflow file you loaded was saved by an older release. The sampler gained widgets over several versions, and when a saved file has a different number of values than the node has dials,

    ComfyUI fills them in order and everything after the mismatch lands on the wrong dial. In that case load H3_Seamless_Chain_v2.json fresh out of the current zip rather than reusing your saved copy - then re-enter any settings you had changed.

    drowai443Aug 11, 2026· 2 reactions
    CivitAI

    You're definitely getting there, man. This is good stuff.

    Your own video above has a cut and isn't seamless though. And the image degradation is pretty severe by the end, like WAN 2.2. Seems like you did fix color shift and audio, and there are no overbright frames at the seams, so this is insane progress for only a few days.

    Keep up the great work.

    joeygambino
    Author
    Aug 12, 2026· 3 reactions

    Yeah, I am working on the degradation, expecting to have some progress by morning. The cut.. I don't even know what happened there, it's been pretty steadily working for me.

    wallmonster151Aug 12, 2026
    CivitAI

    Hey, first as almost everyone else has already said. Thanks for your amazing work! And also for being so engaged on follow ups!

    I managed to get your Riftcast Studio up and running the other day. This flow was looking like a seamless startup for me after I got rid of a Fantasy Talking GGUF node conflict. Then unfortunately the flow ran to about 85% before throwing a math error. I believe I had everything in place as the models auto-populated when I loaded the H3_Seamless_Chain_v2 flow. This was just with the stock images and prompts.

    H3 Multishot Sampler + Memory (long form)

    Error log

    # ComfyUI Error Report ## Error Details - Node ID: 30 - Node Type: H3MultishotMemorySampler - Exception Type: RuntimeError - Exception Message: RuntimeError: mat1 and mat2 shapes cannot be multiplied (3680x1152 and 3456x1152)

    I ran the update patch and it came back success. Apologies if this is one of those long since asked and answered. Feel like I had a pretty solid look around for someone with the same issue and cam up empty.

    Thanks


    LemmingWolf01Aug 12, 2026

    @joeygambino Hi, just to add a little more to this as I've been having the same error. My settings in the master control are 640x960 (2:3) I'm only using one image, or at least it's the only one activated and that's in the Identity Anchor Image Node and the resolution is 1024x1536 (2:3) so from a resolution perspective they should be compatible?

    The things I have noticed:

    1. It only errors if I'm using a GGUF text encoder, safetensors work fine although it doesn't bring the image in as the first image for the shot.

    2. After the error if I look at the parameters for the node, which I assume reflect the state of play when the fatal error occurred, width and height are reported as 764x1344 which isn't 2:3. However, if I do get a successful run (using a safetensor text encoder) the finished video is as specified in the master control i.e. 640x960.

    I don't know what any of this means, it's mostly all well above my brain cell count, but I thought I'd let you know incase it helps.

    Edit: Just to mention after reading wallmonster's latest comments, my error occurs pretty much as soon as it hits the sampler.

    wallmonster151Aug 12, 2026

    @joeygambino Thanks for the quick response!

    Should have clarified but as with Lemming below that was with the gguf models. That first run was with only the place holder 768x768 images in there slots. All use image toggles were turned off which should have defaulted to whatever T2I resolution the workflow was saved at and the example text in place. I do not believe I flipped a single toggle on that run. When I was first starting with comfy I threw a lot of math errors with text encoder mismatches but as you said those where always when the first merge happened. Here it goes pretty much all the way until it is getting ready to move off the ksampler. Rough previews were generating that looked to match your example script. Maybe trying to merge the shots?.

    Anyway I'll keep digging through my settings. I do have a pretty robust local machine so I'll give the safetensors version a try. My main interest in gguf version is iteration speed as I am learning, as we see here it is sometimes better to fail fast. Local storage space is another big plus for gguf. Like most of your users I am a bit of a hoarder and hesitant to delete anything. Either I have happy memories of one good run with that file or it is on my mental list to go back and figure out how to optimize later.

    Edit here: It was not actually at 85%. The ksampler goes to 100% of the first shot and the error throws on the handoff to the second shot. I made sure the image input toggles were all turned off and for extra security bypassed all of their loaders as well as the audio anchor loader. I tried 1152x1152 hoping for a direct match to the model but something, somewhere is adding a little to the image width no matter what I put in the master. I changed the reference image size toggle from match to max also with no joy.

    Edit 2: I thought I kept everything 100% unchanged when I first loaded the workflow but I may have been a little too proactive. When trying to figure out where the extra width is coming from I changed the text encoder sidecar setting from referencing the actual file to auto and at least T2V it was able to join 2 shots and run to completion. It's certainly possible that I populated that field myself on the initial run. Testing I2V now with the sidecar on auto.

    Edit 3: It runs to completion I2V with the sidecar set to auto.

    Early days after only one run but the initial run off of the same reference image and resolution did not seem to generate the same quality as your Riftcast/JoyEcho workflows. Now that I have completed a run I'll move up to the Q8 model and see how that goes.

    Thanks again!

    joeygambino
    Author
    Aug 12, 2026

    Found it - and it is my bug, not your setup. Ignore my earlier answer about resolutions; that was wrong, sorry for the detour.

    It's the mmproj_name widget on the H3 CLIP Loader. Naming a file there went down a different code path than (auto) and skipped the key-renaming step, so the vision tower loaded under names nothing reads. That is why it always died at the shot-2 handoff, why it was GGUF-only, and why nothing you changed about resolution or image toggles helped.

    Manual fix, in order of least effort:

    1. Set mmproj_name back to (auto). If it loads, you are done.

    2. If (auto) then says "No vision sidecar resolved", the pairing is by filename - the mmproj must sit in the same folder as the encoder and contain the encoder's name minus its quant suffix:

    So either rename the mmproj to match, or make it the only file with "mmproj" in the name in that folder - the loader falls back to "if there is exactly one, use it."

    3. If you would rather not touch your model folder, one line in custom_nodes/ComfyUI-H3-Multishot/h3_multishot_utils.py. Find:

    if mmproj_name and mmproj_name != "(auto)":

    and change it to:

    if False and mmproj_name and mmproj_name != "(auto)":

    That makes the widget inert and forces the working path. Restart ComfyUI. It is a workaround, not the fix - the real one keeps the widget working for people with split folders.

    Or just wait. It is already fixed and verified on my side, and the next release is close — it also carries two other things that stop the workflow running for anyone who installed from here: the accelerator nodes shipped switched on (so a clean install could not queue at all), and a widget mismatch that threw "The value 1 for reference_image_size is not available". If you are not blocked today, the update will be the cleaner path.

    Thanks again - @wallmonster151, your Edit 2 is what found this. It would have stayed hidden for a long time otherwise.

    @LemmingWolf01 - the 768x1344 you saw on the node after the error is a display quirk, not the cause: width and height are driven by links from MASTER CONTROLS, so the widget keeps showing its own stored default. Your render really was 640x960. Separately, tell me which safetensors encoder you used when the image did not come in as the first frame and I will chase that one too.

    joeygambino
    Author
    Aug 12, 2026

    Found it - and it is my bug, not your setup. Ignore my earlier answer about resolutions; that was wrong and I am sorry for the detour.

    It is the mmproj_name widget on the H3 CLIP Loader. Naming a file there went down a different code path than (auto) and skipped the key-renaming step, so the vision tower loaded under names nothing reads. That is why it always died at the shot-2 handoff, why it was GGUF-only, and why nothing you changed about resolution or image toggles helped.

    Manual fix, in order of least effort:

    1. Set mmproj_name back to (auto). If it loads, you are done.

    2. If (auto) then says "No vision sidecar resolved", the pairing is by filename - the mmproj must sit in the same folder as the encoder and contain the encoder's name minus its quant suffix:

    MiniMax-H3-encoder-Q5_K_M.gguf + MiniMax-H3-encoder-mmproj-F16.gguf pairs

    MiniMax-H3-encoder-Q5_K_M.gguf + mmproj-F16.gguf does not

    So either rename the mmproj to match, or make it the only file with "mmproj" in the name in that folder - the loader falls back to "if there is exactly one, use it."

    3. If you would rather not touch your model folder, one line in custom_nodes/ComfyUI-H3-Multishot/h3_multishot_utils.py. Find:

    if mmproj_name and mmproj_name != "(auto)":

    and change it to:

    if False and mmproj_name and mmproj_name != "(auto)":

    That makes the widget inert and forces the working path. Restart ComfyUI. It is a workaround, not the fix - the real one keeps the widget working for people with split folders.

    Or just wait. It is already fixed and verified on my side, and the next release is close — it also carries two other things that stop the workflow running for anyone who installed from here: the accelerator nodes shipped switched on (so a clean install could not queue at all), and a widget mismatch that threw "The value 1 for reference_image_size is not available". If you are not blocked today, the update will be the cleaner path.

    Thanks again - @wallmonster151, your Edit 2 is what found this. It would have stayed hidden for a long time otherwise.

    @LemmingWolf01 - the 768x1344 you saw on the node after the error is a display quirk, not the cause: width and height are driven by links from MASTER CONTROLS, so the widget keeps showing its own stored default. Your render really was 640x960. Separately, tell me which safetensors encoder you used when the image did not come in as the first frame and I will chase that one too.

    wallmonster151Aug 12, 2026

    @joeygambino Thanks man! Coincidentally I saw you were updating that file on git when I was digging around. I almost threw in the new utils file to test. I will do that on the next run. In truth I spent a fair amount of time hacking around in that file last night with no joy so I did a full revert to confirm the issue before reaching out.

    LemmingWolf01Aug 13, 2026

    @joeygambino Hi, thanks for the updates and I can confirm that GGUF text encoders work now. With regard to the first frame issue, the safetensors encoder I was using was just the stock "qwen3vl_32b_minimax_h3_nvfp4_awq" I've since tried with a GGUF + mmproj and still no first frame from my Identify anchor image. Perhaps I'm not fully understanding the process and it's something I'm doing wrong. As it stands I have the:

    Reference Gate OFF,

    FFLF Plate Gate OFF,

    Identity Anchor gate ON (with image),

    Continuity = first_frame,

    and using a fl2va model (I have tried a ref2va model as well, though in this instance it shouldn't be required, should it?)

    As I understand things, that should produce a shot with the first frame as per my Identity Anchor image.

    As I said, maybe it's something I'm doing or not doing.

    Thanks for all your fantastic work and help.

    Workflows
    MiniMax H3

    Details

    Downloads
    253
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/11/2026
    Updated
    10/2/2026
    Deleted
    -

    Files

    minimaxH3MultishotSeamlessChain_v20SeamlessChain.zip

    minimaxH3MultishotSeamlessChain_v20SeamlessChain.zip