CivArchive
    MiniMax-H3 Multishot — chained shots, one master, with audio - v1.3
    NSFW
    Preview 138837906

    Support

    Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:

    Three ways to drive MiniMax-H3, and a VRAM fix that makes it usable on a 32 GB card. Chain shots from a script into one long piece; pin keyframes anywhere in a clip; or drive identity from reference images, video and voice. All in one node pack, all producing video and audio.

    Keyframes at any position

    Stock ComfyUI pins H3 keyframes to the first and last frame only and raises only first/last keyframe anchors are supported for anything else.

    That is a positional-maths limit, not a model limit. Both stock cases are the same expression, because sum(_video_t_spans(latent_t)) == FRAME_RESCALE * frame_count:

    cond_t = text_len + FRAME_RESCALE * pixel_index

    — which is defined for every frame, not just the two endpoints. So you can hand H3 up to six anchor images and say when each one should happen, as fractions (0, 0.5, 1) or absolute frame indices.

    Measured on an RTX 5090, 243 frames, one anchor at pixel frame 121: the rendered frame most resembling the anchor image was frame 122 — the requested position, off by one — arrived at by continuous motion with no cut (peak frame-to-frame change 2.3× the median), and the audio ran unbroken straight through it. A three-anchor run at 0 / 0.5 / 1 landed the second and third on frames 121 and 242 exactly.

    The patch is applied in memory — it does not edit any ComfyUI file. It self-tests against the stock formula before committing and rolls itself back if first/last positions do not reproduce exactly, so a future ComfyUI change degrades to “interior anchors unavailable” rather than to broken renders.

    It moves, or it cuts — and your images decide which

    Anchor images with a plausible camera path between them (same place, different angle or framing) make H3 interpolate: a real move that arrives on time. Images with no possible path — a kitchen and a diner — make it cut, then hold.

    That is the model being sensible, not a limitation of the node: stock first/last does exactly the same thing when given such a pair. And the cut case earns its keep, because it is a timed shot change inside a single generation — one generation means one continuous audio stream, so the voice does not get re-derived and there is no seam to hide.

    Saying when: percentages, indices and ranges

    Anchor positions take whichever form suits the shot:

    0%, 50%, 100%     percentages
    0, 121, 242       absolute frame indices
    0-9, 352-361      inclusive ranges
    30%-20%           descending: reverses that section of the batch

    Ranges matter when you have a burst of frames rather than a single anchor — several frames clustered at each end to pin a complex move, or a run kept from a source video. For that there is an images_batch input that takes any number of anchors, and it adds to the six individual slots rather than replacing them.

    Both of those came from @poltergeisha360, who asked for the batch input here on Civitai and then wrote both features and sent pull requests. The percentage syntax also fixes a genuine trap of mine: a bare 1 used to mean the last frame, not frame 1, so addressing an early frame absolutely meant writing 1.0001. Saved workflows keep working — a bare non-integer at or below 1.0 is unambiguous and still reads the old way, while an all-integer 0, 1 takes the new absolute meaning and logs a warning rather than silently anchoring a different frame.

    Models in subfolders show up

    Both loaders scan diffusion_models and text_encoders recursively, so a GGUF filed under gguf/ appears in the dropdown. It did not before, while ComfyUI-GGUF's own loader listed the same file fine — the scan exists because .gguf is not in ComfyUI's supported_pt_extensions, so the normal file list never returns it.

    Which workflow

    • A script of several shots, chained into one piece — H3_Multishot_AIO.json

    • Specific frames at specific times, one continuous take — H3_Keyframes.json

    • 2–5 minutes without identity drifting — H3_Multishot_MEMORY.json

    • Identity from reference images, video or voice — Hard Mode, a separate download

    Reference-to-video has landed, as its own release: MiniMax-H3 Hard Mode — identity from reference images, video and audio instead of a start frame, up to 9 images, 3 videos, 3 soundtracks and 3 standalone audio clips. It uses this pack's nodes, so install this first. On Hugging Face and GitHub. Also on Civitai.

    References and keyframes are mutually exclusive — worth knowing before you wire the two together. This is ComfyUI core behaviour, not a choice made here: model_base.py assigns cond_video_latents for references, discarding keyframe latents while the keyframe layout rows survive — the packed sequence then desyncs into a shape-mismatch crash. There is no “reference images plus start frame” mode. Pick one per shot.

    ~4× faster on 32 GB cards

    The text encoder is evicted before sampling. The Qwen3-VL encoder (~16.5 GB even at Q4) and the H3 DiT (~25 GB) do not co-fit on a 32 GB card, so the DiT was loading partially and streaming ~19 GB from system RAM on every sampling step. If you have ever seen this in your log:

    loaded partially; 6423 MB usable, 5847 MB loaded, 19363 MB offloaded

    that was it. Measured on an RTX 5090 — 960×544, 124 frames, 20 steps, ref2va-Q5_1, one reference image: 12.3 min with the encoder evicted. The un-evicted run of that same render was killed past 90 min without finishing, so there is no honest completion time to quote against it. Every sampler in this pack evicts, including the keyframes node, and prints TE evicted; NN.N GB free for the DiT.

    Treat this as a cliff, not a curve. The encoder (~16.5 GB at Q4) and the DiT (~25 GB) do not co-fit in 32 GB, so either the DiT is resident and you get normal speed, or it streams ~19 GB every sampling step. The size of the slowdown depends on how far over your card you are, not on resolution or frame count directly — one speed-up ratio would not generalise, so there is not one here.

    Quick fixes — read this first

    • Red / missing nodes when a workflow loads
      Why: pack not installed, or ComfyUI not restarted.
      Fix: install ComfyUI-H3-Multishot (Manager > Install via Git URL), restart, then hard-refresh the browser tab — the frontend caches node definitions.

    • GGUF errors with “unknown model architecture”
      Why: ComfyUI-GGUF does not know MiniMax-H3 out of the box.
      Fix: run python apply_gguf_arch_patch.py from the pack folder (one line, idempotent), restart. This is for the DiT only — the text encoder is Qwen3-VL and needs no patch.
      Only want the DiT working? The patch is also downloadable on its own (2 KB) if you are running ComfyUI's built-in MiniMax-H3 nodes with the safetensors encoder and do not need this pack at all. It is included here, so installing this pack is enough.

    • Reference audio crashes the sampler with a shape mismatch
      Why: your clip is mono. The audio VAE encodes [B, 2, L] and the layout reserves exactly two channels, so a mono reference produces half the rows it reserved and dies deep inside the model with no useful message. Nothing in stock converts it.
      Fix: the H3 Reference Audio node — forces stereo 32 kHz and trims length. It is already wired in the hard-mode workflow.

    • The model ignores my reference image
      Why: reference blocks are labelled in the prompt and the numbering is 1-based while the input slots are 0-basedref_image_0 is <Picture 1>. If you never name it in the text, the model has no reason to bind it.
      Fix: write <Picture 1> is the woman. <Picture 2> is the room. Note a reference video with a soundtrack consumes an <Audio j> ordinal before your standalone clips.

    • Multi-shot ignores my reference image entirely
      Why: missing mmproj. It is required for chaining, not just for reference images — chaining feeds the previous shot's last frame through the encoder's vision path.
      Fix: download the -mmproj file alongside the encoder and keep both filenames exactly as downloaded, in the same folder. The loader pairs them by name.

    • Length change errors out
      Why: H3 hard constraint — frame counts live on a 17k+5 grid.
      Fix: 226, 243, 260… the widget steps by 17 so it keeps you legal. 243 ≈ 10s; 362 ≈ 15s is the trained ceiling.

    • Speech turns to gibberish
      Why: usually an under-filled shot, not an over-long one. Speech runs about 2.5 words/second, so a 243-frame shot wants roughly 22–25 spoken words; give it eight and the model invents sound to fill the dead air.
      Fix: match dialogue length to shot length, and if you want silence, script it (“she listens, saying nothing”).

    • An object morphs into something else at a seam
      Why: chaining hands each shot the previous final frame. If shot 1 ends on the cameraman holding his camcorder and shot 2 is filmed FROM that camcorder, the model must explain a device in a hand that should not be in frame — so it invents one.
      Fix: end every shot on what the NEXT shot expects to see.

    • Out of VRAM, or renders crawl
      Why: 33B of weights — and see the eviction section above.
      Fix: use the GGUFs, Q5_1 for 24–32 GB, Q4_0 for 16 GB. The file does not need to fit in VRAM; ComfyUI streams the overflow. Expect ~10 min per 10s shot on a 5090-class card.

    Writing a script — this is most of the quality

    The identity lock is description density, not assertion. Every shot is an independent conditioning pass: the model rebuilds the person from your text each time. Writing “the same woman, same face, same wardrobe” asserts continuity without supplying what is needed to rebuild it, and the face drifts. Re-describing 6–8 concrete attributes verbatim in every shot is what actually holds it:

    She is an attractive American woman in her mid twenties with warm hazel
    eyes, a friendly confident smile, light freckles, shoulder-length auburn
    hair tucked behind one ear, small gold stud earrings, and a relaxed
    sage-green blouse. Her voice is a clear warm young woman's voice in a
    casual American accent.

    Do the same for the voice: one short concrete line, repeated verbatim. Flowing prose beats SHOT: / Audio: labels.

    What you need

    A word on expectations

    MiniMax-H3 is a 33B joint audio+video model and this pack started days after the weights landed. It works, and the measurements on this page are from real renders on one consumer GPU — but you may still need to tune to YOUR machine. If you get stuck, comment here or open a GitHub issue. I answer.

    Everything else I've published

    Support

    Everything I publish is free and stays free. If it saved you a night of debugging, tips keep the 5090 warm: Ko-fi · GitHub Sponsors · Liberapay.

    Description

    Keyframe positions get percentages and ranges, keyframes take an image batch, and the model loaders stop hiding your GGUFs.

    Keyframes: image batches, percentages and ranges

    Two contributions from @poltergeisha360, who asked for the first one here on Civitai and then wrote both himself and sent pull requests (as @viralesveras on GitHub).

    An images_batch input. Six individual slots runs out quickly if you want several frames clustered at each end to pin complex motion, or a set of frames kept from a source video. The batch takes any number of anchors. Merged with one change: it adds to the six individual slots rather than replacing them — as originally written, wiring image_1 and then adding a batch dropped image_1 silently.

    Percentages and inclusive ranges in positions. This fixes a genuine trap of mine: a bare 1 meant the last frame, not frame 1, so addressing an early frame absolutely meant writing 1.0001. That is how the ambiguity got found — in real use.

    0%, 50%, 100%     percentages
    0, 121, 242       absolute frame indices
    0-9, 352-361      inclusive ranges
    30%-20%           descending: reverses that section of the batch

    Your existing workflows keep working. A bare non-integer at or below 1.0 is unambiguous — nobody means "frame index 0.5" — so a saved 0, 0.5, 1 is still read the old way, and logs the percentage spelling to switch to. An all-integer 0, 1 is genuinely ambiguous, so it takes the new absolute meaning and warns, rather than silently anchoring a different frame than it used to.

    Loaders find GGUFs in subfolders

    Both model loaders scanned only the top level of diffusion_models and text_encoders, so anything kept in a gguf/ subfolder was invisible in the dropdown — while ComfyUI-GGUF's own loader listed the same files fine. They walk the tree now.

    If your models showed up on one machine and "disappeared" on another, this was why: the scan exists because .gguf is not in ComfyUI's supported_pt_extensions, so the normal file list never returns it. It just was not recursive.

    New models: curve-form GGUFs, ~40% smaller

    Not part of this zip, but worth knowing about — there is now a second DiT repo carrying -curve- files built from MiniMax's pruned checkpoints. About 40% of this model was never unique data: each block carried a 96768×2688 modulation matrix, and those turn out to be a smooth function of the timestep rather than 51 independent tensors.

    TierCurve formOriginalQ8_021.5 GBnever built — too large to be worth itQ5_115.2 GB25.9 GBQ4_011.5 GB19.9 GB

    That Q8_0 is smaller than the Q5_1 that has been shipping since day one, and it fits a 24 GB card. Verified by rendering, same graph and same seed, changing only the model file: swapping to curve form moved the output about a third as far as dropping one quant tier does, with per-frame motion in family. Files and the full comparison table are on Hugging Face — a separate repo, so the original-form quants stay where they are for anyone on an older ComfyUI or who prefers the larger files.

    Curve files need ComfyUI 0.30.0+. On older builds they will not load at all, because the shape of the modulation weights changed. The original-form files remain up for anyone pinned to an older version.

    Full contents

    • H3_Multishot_AIO.json — easy mode: one script to many chained shots

    • H3_Multishot_MEMORY.json — 2–5 minute pieces with an identity anchor

    • H3_Keyframes.json — anchors at any position, single pass, unbroken audio

    Every one of those three has a render on disk that its own node produced, verified from the render's embedded graph rather than from the file existing.

    FAQ

    Comments (10)

    sdktertiaire2Aug 5, 2026· 1 reaction
    CivitAI

    hello and thanks for your work. You wrote " RTX 5090: ~60 min → ~15 min ". For wich width/ height and laps of time/frames ?

    joeygambino
    Author
    Aug 5, 2026

    Good catch - that line's missing its conditions, and now that I've gone back to the logs it's also just not a good pair of numbers.

    The real run, pulled from the render's own embedded workflow:

    960 x 544, 124 frames (5.2 s at 24 fps), 20 steps

    ref2va-Q5_1, one reference image, RTX 5090, --reserve-vram 7

    -> 12.3 minutes with the encoder evicted

    The un-evicted version I killed somewhere past 90 minutes, so it never finished. That means "~60 min" is a number I don't actually have, and "~15 min" was 12.3 rounded up. I'll fix the model page - thanks for asking instead of trusting it. A properly controlled pair, where both sides ran to completion or a deliberate cut-off (480x864, 124 frames, one reference VIDEO):

    no eviction: killed at 46 min, 98% "utilization" at ~172 W

    eviction on: 6.5 and 7.7 min on two clean runs, 355-390 W

    The thing that matters more than any ratio: this is a cliff, not a curve. The Qwen3-VL encoder is ~16.5 GB even at Q4 and the DiT is ~25 GB, so on a 32 GB card they just don't co-fit. Either the DiT is resident and you get normal speed, or it streams ~19 GB from system RAM every sampling step. How bad it gets depends on how far over your card you are - which is why quoting one speed-up ratio was misleading of me regardless of the numbers.

    Two tells. In your log:

    loaded partially; 6423 MB usable, 5847 MB loaded, 19363 MB offloaded

    and on the card, watch POWER DRAW, not utilization. A GPU thrashing weights between RAM and VRAM still reports ~98% while pulling a fraction of its rated watts. 146-172 W on a 5090 is thrashing; 400 W+ is real work. I chased the utilization number for hours before the wattage told me what was going on. If you're on 24 GB, the curve-form Q8_0 is worth a look - 20 GiB resident on a 3090 with room to spare, 6.8 min for 123 frames, and it's smaller and faster than the original-form Q4_0.

    ugurdoyduk341Aug 5, 2026· 1 reaction
    CivitAI

    Is it just me that face, eye and lip artifacts have decreased with 1.3? Anyway, great update.

    joeygambino
    Author
    Aug 5, 2026

    Thank you!

    JustTrying2026Aug 5, 2026· 1 reaction
    CivitAI

    Seen someone mention in discord that heunpp works really well for better audio. But I don't see a way to change it in this node. Have you tried this heunpp?

    joeygambino
    Author
    Aug 5, 2026

    I will work on exposing the sampler and scheduler in the AIO node for the next version, likely have it up tomorrow.

    aikoyamaAug 6, 2026· 1 reaction

    My experience with heunpp2 has been very positive for audio but it is slow. Very slow. Maybe twice as slow as res_multistep.

    joeygambino
    Author
    Aug 6, 2026

    @JustTrying2026 v1.4 uploaded with sampler and scheduler exposed

    shiftycheshireAug 6, 2026· 1 reaction
    CivitAI

    Have you done any testing for taking an existing video and extending it? MiniMax H3 has the capabilities to do so but from my own testing is oftne has issues with audio or color grading. I was wondering if a similar method coould be applied there as well (as in, take an existing video + prompt as as the first 'shot' instead of just an image + prompt)

    joeygambino
    Author
    Aug 6, 2026

    I have not yet, but I will put it on the list for next workflows!

    Workflows
    MiniMax H3

    Details

    Downloads
    243
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/5/2026
    Updated
    8/7/2026
    Deleted
    -