CivArchive
    Minimax H3 LONG VIDEO (FL2VA, REF2VA) - FL2VA + REF2VA V4.0
    NSFW
    Preview 142021334
    Preview 142021349

    MiniMax H3 Long Video Workflows — V4.0

    Support the project

    If this workflow helped you, you can support future updates using the Tip button on my Civitai profile or make an optional USDT donation.

    Token: USDT

    Network: TRON (TRC20)

    Address: THLU3yu8VpoPqP4bMLCp86FRhoCweJ2Qz4

    Please send only USDT through the TRON (TRC20) network. Assets sent through another network may be permanently lost. Thank you!

    V4.0 is a complete redesign of the MiniMax H3 long-video workflows. It is smaller, easier to configure, and supports from 1 to 12 sequential sections.

    The workflow transfers protected video and audio latent context between sections, allowing MiniMax H3 to generate a longer continuous video instead of unrelated clips.

    Included workflows

    • FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.

    • REF2VA — supports image, video, video-soundtrack, and standalone audio references.

    This release contains two archives:

    • Workflows + Instruction Prompts — both V4.0 workflows and bilingual Russian/English prompt guides.

    • h3_direct_latent_head — the shared custom node pack required by both workflows.

    Download and install both files.

    What changed since V3.0

    • The large fixed graph was replaced with compact dynamic generation.

    • The number of sections can be selected from 1 to 12.

    • You mainly configure only the MASTER and H3 MODEL STACK nodes.

    • The learned latent upscaler and HQ refinement were removed.

    • Every section is generated directly at the selected native resolution.

    • Only one selected model branch and one attention backend are used.

    • Turbo 4-step, Turbo 8-step, Native 20-step, and Native 25-step profiles are available.

    • Seed increment per section is enabled by default to reduce repetitive motion.

    • FL2VA and REF2VA use separate validated prompt formats.

    • REF2VA now supports image, video, video-audio, and standalone audio references.

    • REF2VA includes a safer reference policy for long-video continuation.

    • Final audio is decoded once from the complete cumulative AV timeline.

    • Model Preview Override is included and enabled.

    V4.0 prioritizes native generation, continuity, and a simpler workflow over experimental upscaling and refinement.

    How the long-video system works

    Each section is a generation window. A protected part of the previous section’s video and audio latent is transferred into the next section.

    The recommended overlap is 22 frames. It helps preserve:

    • the current pose and movement;

    • character appearance;

    • camera position and direction;

    • lighting and scene geometry;

    • continuous sound and ambience.

    With the default settings:

    • 6 sections;

    • 141 frames per section;

    • 22 protected context frames;

    • 24 FPS;

    the final video contains 736 frames, or approximately 30.67 seconds.

    The first section contributes all 141 frames. Each following section contributes 119 new frames because 22 frames are reused as continuity context.

    Installation

    1. Stop ComfyUI completely.

    2. Delete the previous folder:

      ComfyUI/custom_nodes/h3_direct_latent_head

    3. Extract the new h3_direct_latent_head folder into:

      ComfyUI/custom_nodes/

    4. Do not merge files from different node versions.

    5. Install the required external custom nodes.

    6. Place the MiniMax H3 models, text encoder, video VAE, audio VAE, and optional LoRAs in the standard ComfyUI model folders.

    7. Restart ComfyUI.

    8. Open the required FL2VA or REF2VA V4.0 workflow.

    H3 MODEL STACK uses portable AUTO selections by default. If several matching files are found, select the correct model or LoRA manually.

    FL2VA modes

    Control the generation mode using the START and END image inputs:

    • T2V: START and END are bypassed.

    • I2V: enable START only.

    • Last-frame guidance: enable END only.

    • FL2V: enable both START and END.

    Each active section requires a separate prompt with three blocks:

    • integrated_multimodal_description

    • overall_soundscape

    • non_diegetic_music

    S1 establishes the subject, scene, camera, movement, and sound. S2 and later prompts should describe the next stage of the same continuous action.

    Do not repeat the same action as a new command in every section, because the model may replay it repeatedly.

    A complete Russian/English FL2VA prompt guide is included.

    REF2VA references

    REF2VA supports:

    • up to 4 image references;

    • up to 3 video references;

    • soundtracks from reference videos;

    • up to 3 standalone audio references.

    Enable inputs consecutively inside each type: input 1 before input 2, and input 2 before input 3.

    Reference videos are loaded at 24 FPS and limited to the first 360 frames / 15 seconds.

    For the first test, use:

    • one reference;

    • 0.31 MP resolution;

    • 22F context;

    • 0F audio history;

    • reference_policy = S1 ONLY.

    Real image, video, and audio references are presented to the model only during S1. S2–S12 continue through the protected AV latent.

    This reduces the risk of later sections suddenly reproducing:

    • the original reference image;

    • the original pose or composition;

    • the reference background;

    • the original video movement;

    • the reference audio as a restarted sound.

    Use <Picture N>, <Video N>, and <Audio N> tags only in the S1 prompt. Do not use reference tags in S2–S12.

    EVERY SECTION

    This mode presents all references again during every section.

    It can provide stronger reference adherence, but it may also reproduce the original pose, scene, motion, background, or audio as a separate insert. It also requires more processing time and memory.

    Each active REF2VA section requires a separate prompt with six blocks:

    • subject_definitions

    • summary

    • retention_analysis

    • detailed_description

    • overall_soundscape

    • non_diegetic_music

    A complete Russian/English REF2VA prompt guide with S1 and continuation templates is included.

    • Resolution: 0.31 MP

    • Section length: 141F

    • Context: 22F balanced

    • Audio history: 0F off

    • Seed strategy: increment per section

    • Reference image size: MATCH

    • REF2VA reference policy: S1 ONLY

    • Video shift: 12

    • Audio shift: 3

    Increase the resolution and number of references only after a successful test.

    Sampling and LoRAs

    Available sampling profiles:

    • Turbo 4 steps

    • Turbo 8 steps

    • Native 20 steps

    • Native 25 steps

    The sampling profile must match the selected LoRA. For example, use a 4-step Turbo LoRA with the 4-step profile and an 8-step Turbo LoRA with the 8-step profile.

    FL2VA and REF2VA require LoRAs intended for the corresponding model type.

    Attention backends

    The workflow uses exactly one selected attention backend:

    • COMFY KITCHEN

    • SAGE

    • PYTORCH

    Comfy Kitchen is selected in the supplied workflows. If it is unavailable in your ComfyUI build, select SAGE or PYTORCH.

    SAGE requires ComfyUI-KJNodes.

    Required models

    MiniMax H3 models:

    Comfy-Org/MiniMax-H3

    Compatible Heretic Text Encoder — choose one:

    Optional LoRA sources:

    Required custom nodes

    • ComfyUI-KJNodes — required only for SAGE Attention.

    • ComfyUI-VideoHelperSuite — required for REF2VA video references.

    • h3_direct_latent_head — included with this release and required by both workflows.

    No latent-upscaler node is required.

    Model Preview Override and stability

    Model Preview Override is included and enabled so intermediate sampling previews can be monitored during generation.

    On some Windows and PyTorch configurations, preview decoding may occasionally cause:

    Fatal Python error: Aborted

    If this happens:

    1. Refresh the ComfyUI browser page.

    2. Bypass the Model Preview Override node.

    3. Queue the workflow again.

    Bypassing Model Preview Override disables only the intermediate preview. It does not change the final generation, prompts, video continuity, audio timeline, or output quality.

    Comfy model compiler graph breaks is diagnostic information and does not mean that generation has failed.

    The node pack and workflow contracts were tested against ComfyUI 0.34.0.

    Optional final post-processing

    If additional upscaling or frame interpolation is required, process the completed video afterward with:

    DLSS 5 Visual Enhancer

    It supports RTX Video Super Resolution, DLSS Neural Rendering, and optional frame generation. It is a separate application and is not required for these ComfyUI workflows.

    Download DLSS 5 Visual Enhancer

    MiniMax H3 Long Video Workflows — V3.0

    V3.0 includes two six-section MiniMax H3 long-video workflows and the h3_direct_latent_head custom node.

    • FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.

    • REF2VA — supports up to four optional image references across S1–S6 and HQ refinement.

    Main features

    • Direct transfer of video and audio latents between sections.

    • Exact phase-aligned 22-frame protected AV context.

    • Accurate 24 FPS video / 40 Hz audio timeline.

    • Optional audio history: 0F, 24F, 48F, or 72F. Recommended: 48F / 2 seconds.

    • Final audio is assembled cumulatively, decoded once, and trimmed to the exact video duration.

    • Automatic section-length and final-duration calculation.

    • H3 denoise-mask compatibility fix based on ComfyUI PR #15988.

    • MASTER node for resolution, aspect ratio, quality, continuity, section length, audio history, and S1–S6 prompts.

    • Individual section checkpoints and final assembled video.

    • Separate FL2VA and REF2VA prompt-writing guides.

    Quality modes

    • NATIVE — generation without latent upscaling or HQ refinement.

    • DIRECT — learned latent upscale without diffusion redraw. Recommended for maximum stability.

    • HQ REFINE — learned latent upscale followed by three-step high-resolution refinement. This can improve detail but may slightly redraw some frames.

    Upscaling and HQ refinement do not regenerate the audio.

    Turbo LoRA affects S1–S6 and HQ refinement. The secondary motion/style LoRA affects only S1–S6.

    Video length

    Valid section lengths follow this grid:

    section_frames = 5 + 17 × N

    Minimum value: 39.

    For six sections:

    final_frames = 6 × section_frames − 110

    The first section keeps its full length. Each following section contributes section_frames − 22 new frames.

    Models

    MiniMax H3 models:

    Comfy-Org/MiniMax-H3

    Heretic Text Encoder — choose one:

    LoRA sources:

    Learned latent upscaler:

    Required custom nodes

    The external latent-upscaler node must be installed for the workflows to load, even when using NATIVE mode.

    Installation

    1. Stop ComfyUI.

    2. Delete the previous ComfyUI/custom_nodes/h3_direct_latent_head folder.

    3. Copy h3_direct_latent_head from the V3.0 archive into ComfyUI/custom_nodes/.

    4. Install the other required custom nodes and models.

    5. Restart ComfyUI and open the FL2VA or REF2VA workflow.

    Do not merge files from different versions of the custom node.

    Model paths stored in the workflows are examples. Select the corresponding files installed on your system if the paths or filenames differ.

    Usage

    Configure the workflow through the MASTER node:

    • quality mode and resolution;

    • aspect ratio;

    • continuity profile;

    • section_frames;

    • audio-history length;

    • prompts for S1–S6.

    FL2VA modes:

    • T2V: keep START FRAME and END FRAME bypassed.

    • I2V: enable START FRAME only.

    • First/last-frame generation: enable both START FRAME and END FRAME.

    In REF2VA, all four references are bypassed by default. Enable them consecutively: REF 1, REF 2, REF 3, and REF 4. Use only the corresponding <Picture 1><Picture 4> tags in the prompts.

    For the first REF2VA test, use REF 1 with the DIRECT quality mode.

    Recommended final post-processing

    After assembling the final video, I recommend processing it with DLSS 5 Visual Enhancer.

    It can improve perceived detail, upscale the completed video, and optionally increase the frame rate using DLSS Frame Generation.

    DLSS 5 Visual Enhancer is optional and runs separately from ComfyUI.

    Download DLSS 5 Visual Enhancer

    My Telegram channel:

    https://t.me/aibobpublic

    MiniMax H3 Long Video Workflows — V2.0

    V2.0 includes two six-section long-video workflows:

    • FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.

    • REF2VA — supports up to four optional image references connected across all six sections and HQ refinement.

    Main features

    • Direct transfer of video and audio latents between sections.

    • 22-frame AV context for seamless continuation.

    • Accurate cumulative audio timeline.

    • Automatic section length and final video duration calculation.

    • MASTER node for aspect ratio, resolution, quality preset, section length, and S1–S6 prompts.

    • Native generation or optional learned latent upscale with 3-step HQ refinement.

    • Individual section checkpoints and final assembled video.

    • Compatible with Sage Attention, Fused Modulation, and FFN chunking.

    • Turbo LoRA affects S1–S6 and HQ refinement; the secondary motion/style LoRA affects only S1–S6.

    • Separate FL2VA and REF2VA prompt-writing guides are included.

    Models

    MiniMax H3 models:

    Comfy-Org/MiniMax-H3

    Heretic Text Encoder — choose one:

    Turbo LoRA options:

    Optional latent upscaler:

    Required custom nodes

    Copy h3_direct_latent_head into:

    ComfyUI/custom_nodes/

    Then restart ComfyUI and open the required FL2VA or REF2VA JSON workflow.

    Model and LoRA paths stored in the workflow are examples. Select the corresponding files installed on your system if your folder structure or filenames differ.

    Usage notes

    Edit the MASTER node to select the quality preset, aspect ratio, section_frames, and prompts for S1–S6.

    In REF2VA, all four references are bypassed by default. Enable references consecutively starting from REF 1. The included example prompts assume that REF 1 is active.

    For REF2VA prompts:

    • S1: define references and subjects, identity retention, action, camera, and sound.

    • S2–S6: repeat subject definitions and retention rules, then continue from the transferred latent context without resetting the pose or scene.

    See the included FL2VA and REF2VA prompt guides for complete templates and examples.

    My Telegram channel:

    https://t.me/aibobpublic

    V1.0
    All Minimax H3 models download from:

    https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main

    Heretic Text Encoder int8:

    https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot/tree/main

    Heretic Text Encoder nvfp4:

    https://huggingface.co/sakamakismile/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4/tree/main

    Turbo loras:

    1) https://huggingface.co/DarkRomeo88/MiniMax-H3-turbo-lora-comfyui/tree/main (without larryvrh nodes) and https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main (with larryvrh nodes).

    2) Also recommend turbo loras:

    https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras

    3) And:

    https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors

    Required custom nodes:

    KJNodes:

    https://github.com/kijai/ComfyUI-KJNodes

    ComfyUI-sol-attn:

    https://github.com/Saganaki22/ComfyUI-sol-attn

    I also developed a custom node for transferring encoded latents between sections. This allows the encoded video and audio latents from the previous section to be passed directly to the next section, helping to avoid overcooking during continuation.

    For the workflow to work, download the custom node archive and place the folder from the archive into the ComfyUI/custom_nodes/ folder.

    You can download it from the drop-down list on this page. There is a separate archive for the workflows and a separate archive for the custom node.

    That is, REF2VA formula for all long-video

    S1:

    REFERENCE BINDING + IDENTITY LOCK + STYLE + ACTION + CAMERA

    S2–S6:

    CONTINUE FROM LATENT CONTEXT + IDENTITY LOCK FROM OPENING FRAMES + ENVIRONMENT CONTINUITY + ACTION + CAMERA

    My TG Channel:

    https://t.me/aibobpublic

    Description

    MiniMax H3 Long Video Workflows — V4.0

    V4.0 is a complete redesign of the MiniMax H3 long-video workflows. It is smaller, easier to configure, and supports from 1 to 12 sequential sections.

    The workflow transfers protected video and audio latent context between sections, allowing MiniMax H3 to generate a longer continuous video instead of unrelated clips.

    Included workflows

    • FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.

    • REF2VA — supports image, video, video-soundtrack, and standalone audio references.

    This release contains two archives:

    • Workflows + Instruction Prompts — both V4.0 workflows and bilingual Russian/English prompt guides.

    • h3_direct_latent_head — the shared custom node pack required by both workflows.

    Download and install both files.

    What changed since V3.0

    • The large fixed graph was replaced with compact dynamic generation.

    • The number of sections can be selected from 1 to 12.

    • You mainly configure only the MASTER and H3 MODEL STACK nodes.

    • The learned latent upscaler and HQ refinement were removed.

    • Every section is generated directly at the selected native resolution.

    • Only one selected model branch and one attention backend are used.

    • Turbo 4-step, Turbo 8-step, Native 20-step, and Native 25-step profiles are available.

    • Seed increment per section is enabled by default to reduce repetitive motion.

    • FL2VA and REF2VA use separate validated prompt formats.

    • REF2VA now supports image, video, video-audio, and standalone audio references.

    • REF2VA includes a safer reference policy for long-video continuation.

    • Final audio is decoded once from the complete cumulative AV timeline.

    • Model Preview Override is included and enabled.

    V4.0 prioritizes native generation, continuity, and a simpler workflow over experimental upscaling and refinement.

    How the long-video system works

    Each section is a generation window. A protected part of the previous section’s video and audio latent is transferred into the next section.

    The recommended overlap is 22 frames. It helps preserve:

    • the current pose and movement;

    • character appearance;

    • camera position and direction;

    • lighting and scene geometry;

    • continuous sound and ambience.

    With the default settings:

    • 6 sections;

    • 141 frames per section;

    • 22 protected context frames;

    • 24 FPS;

    the final video contains 736 frames, or approximately 30.67 seconds.

    The first section contributes all 141 frames. Each following section contributes 119 new frames because 22 frames are reused as continuity context.

    Installation

    1. Stop ComfyUI completely.

    2. Delete the previous folder:

      ComfyUI/custom_nodes/h3_direct_latent_head

    3. Extract the new h3_direct_latent_head folder into:

      ComfyUI/custom_nodes/

    4. Do not merge files from different node versions.

    5. Install the required external custom nodes.

    6. Place the MiniMax H3 models, text encoder, video VAE, audio VAE, and optional LoRAs in the standard ComfyUI model folders.

    7. Restart ComfyUI.

    8. Open the required FL2VA or REF2VA V4.0 workflow.

    H3 MODEL STACK uses portable AUTO selections by default. If several matching files are found, select the correct model or LoRA manually.

    FL2VA modes

    Control the generation mode using the START and END image inputs:

    • T2V: START and END are bypassed.

    • I2V: enable START only.

    • Last-frame guidance: enable END only.

    • FL2V: enable both START and END.

    Each active section requires a separate prompt with three blocks:

    • integrated_multimodal_description

    • overall_soundscape

    • non_diegetic_music

    S1 establishes the subject, scene, camera, movement, and sound. S2 and later prompts should describe the next stage of the same continuous action.

    Do not repeat the same action as a new command in every section, because the model may replay it repeatedly.

    A complete Russian/English FL2VA prompt guide is included.

    REF2VA references

    REF2VA supports:

    • up to 4 image references;

    • up to 3 video references;

    • soundtracks from reference videos;

    • up to 3 standalone audio references.

    Enable inputs consecutively inside each type: input 1 before input 2, and input 2 before input 3.

    Reference videos are loaded at 24 FPS and limited to the first 360 frames / 15 seconds.

    For the first test, use:

    • one reference;

    • 0.31 MP resolution;

    • 22F context;

    • 0F audio history;

    • reference_policy = S1 ONLY.

    S1 ONLY — recommended

    Real image, video, and audio references are presented to the model only during S1. S2–S12 continue through the protected AV latent.

    This reduces the risk of later sections suddenly reproducing:

    • the original reference image;

    • the original pose or composition;

    • the reference background;

    • the original video movement;

    • the reference audio as a restarted sound.

    Use <Picture N>, <Video N>, and <Audio N> tags only in the S1 prompt. Do not use reference tags in S2–S12.

    EVERY SECTION

    This mode presents all references again during every section.

    It can provide stronger reference adherence, but it may also reproduce the original pose, scene, motion, background, or audio as a separate insert. It also requires more processing time and memory.

    Each active REF2VA section requires a separate prompt with six blocks:

    • subject_definitions

    • summary

    • retention_analysis

    • detailed_description

    • overall_soundscape

    • non_diegetic_music

    A complete Russian/English REF2VA prompt guide with S1 and continuation templates is included.

    Recommended starting settings

    • Resolution: 0.31 MP

    • Section length: 141F

    • Context: 22F balanced

    • Audio history: 0F off

    • Seed strategy: increment per section

    • Reference image size: MATCH

    • REF2VA reference policy: S1 ONLY

    • Video shift: 12

    • Audio shift: 3

    Increase the resolution and number of references only after a successful test.

    Sampling and LoRAs

    Available sampling profiles:

    • Turbo 4 steps

    • Turbo 8 steps

    • Native 20 steps

    • Native 25 steps

    The sampling profile must match the selected LoRA. For example, use a 4-step Turbo LoRA with the 4-step profile and an 8-step Turbo LoRA with the 8-step profile.

    FL2VA and REF2VA require LoRAs intended for the corresponding model type.

    Attention backends

    The workflow uses exactly one selected attention backend:

    • COMFY KITCHEN

    • SAGE

    • PYTORCH

    Comfy Kitchen is selected in the supplied workflows. If it is unavailable in your ComfyUI build, select SAGE or PYTORCH.

    SAGE requires ComfyUI-KJNodes.

    Required models

    MiniMax H3 models:

    Comfy-Org/MiniMax-H3

    Compatible Heretic Text Encoder — choose one:

    Optional LoRA sources:

    Required custom nodes

    • ComfyUI-KJNodes — required only for SAGE Attention.

    • ComfyUI-VideoHelperSuite — required for REF2VA video references.

    • h3_direct_latent_head — included with this release and required by both workflows.

    No latent-upscaler node is required.

    Model Preview Override and stability

    Model Preview Override is included and enabled so intermediate sampling previews can be monitored during generation.

    On some Windows and PyTorch configurations, preview decoding may occasionally cause:

    Fatal Python error: Aborted

    If this happens:

    1. Refresh the ComfyUI browser page.

    2. Bypass the Model Preview Override node.

    3. Queue the workflow again.

    Bypassing Model Preview Override disables only the intermediate preview. It does not change the final generation, prompts, video continuity, audio timeline, or output quality.

    Comfy model compiler graph breaks is diagnostic information and does not mean that generation has failed.

    The node pack and workflow contracts were tested against ComfyUI 0.34.0.

    Optional final post-processing

    If additional upscaling or frame interpolation is required, process the completed video afterward with:

    DLSS 5 Visual Enhancer

    It supports RTX Video Super Resolution, DLSS Neural Rendering, and optional frame generation. It is a separate application and is not required for these ComfyUI workflows.

    Download DLSS 5 Visual Enhancer

    FAQ

    Workflows
    MiniMax H3

    Details

    Downloads
    159
    Platform
    CivitAI
    Platform Status
    Available
    Created
    9/6/2026
    Updated
    9/7/2026
    Deleted
    -

    Files

    minimaxH3LONGVIDEOFL2VA_fl2vaREF2VAV40.zip

    minimaxH3LONGVIDEOFL2VA_fl2vaREF2VAV40.zip