MiniMax H3 Long Video Workflows — V4.0
Support the project
If this workflow helped you, you can support future updates using the Tip button on my Civitai profile or make an optional USDT donation.
Token: USDT
Network: TRON (TRC20)
Address: THLU3yu8VpoPqP4bMLCp86FRhoCweJ2Qz4
Please send only USDT through the TRON (TRC20) network. Assets sent through another network may be permanently lost. Thank you!
V4.0 is a complete redesign of the MiniMax H3 long-video workflows. It is smaller, easier to configure, and supports from 1 to 12 sequential sections.
The workflow transfers protected video and audio latent context between sections, allowing MiniMax H3 to generate a longer continuous video instead of unrelated clips.
Included workflows
FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.
REF2VA — supports image, video, video-soundtrack, and standalone audio references.
This release contains two archives:
Workflows + Instruction Prompts — both V4.0 workflows and bilingual Russian/English prompt guides.
h3_direct_latent_head — the shared custom node pack required by both workflows.
Download and install both files.
What changed since V3.0
The large fixed graph was replaced with compact dynamic generation.
The number of sections can be selected from 1 to 12.
You mainly configure only the MASTER and H3 MODEL STACK nodes.
The learned latent upscaler and HQ refinement were removed.
Every section is generated directly at the selected native resolution.
Only one selected model branch and one attention backend are used.
Turbo 4-step, Turbo 8-step, Native 20-step, and Native 25-step profiles are available.
Seed increment per section is enabled by default to reduce repetitive motion.
FL2VA and REF2VA use separate validated prompt formats.
REF2VA now supports image, video, video-audio, and standalone audio references.
REF2VA includes a safer reference policy for long-video continuation.
Final audio is decoded once from the complete cumulative AV timeline.
Model Preview Override is included and enabled.
V4.0 prioritizes native generation, continuity, and a simpler workflow over experimental upscaling and refinement.
How the long-video system works
Each section is a generation window. A protected part of the previous section’s video and audio latent is transferred into the next section.
The recommended overlap is 22 frames. It helps preserve:
the current pose and movement;
character appearance;
camera position and direction;
lighting and scene geometry;
continuous sound and ambience.
With the default settings:
6 sections;
141 frames per section;
22 protected context frames;
24 FPS;
the final video contains 736 frames, or approximately 30.67 seconds.
The first section contributes all 141 frames. Each following section contributes 119 new frames because 22 frames are reused as continuity context.
Installation
Stop ComfyUI completely.
Delete the previous folder:
ComfyUI/custom_nodes/h3_direct_latent_headExtract the new
h3_direct_latent_headfolder into:ComfyUI/custom_nodes/Do not merge files from different node versions.
Install the required external custom nodes.
Place the MiniMax H3 models, text encoder, video VAE, audio VAE, and optional LoRAs in the standard ComfyUI model folders.
Restart ComfyUI.
Open the required FL2VA or REF2VA V4.0 workflow.
H3 MODEL STACK uses portable AUTO selections by default. If several matching files are found, select the correct model or LoRA manually.
FL2VA modes
Control the generation mode using the START and END image inputs:
T2V: START and END are bypassed.
I2V: enable START only.
Last-frame guidance: enable END only.
FL2V: enable both START and END.
Each active section requires a separate prompt with three blocks:
integrated_multimodal_descriptionoverall_soundscapenon_diegetic_music
S1 establishes the subject, scene, camera, movement, and sound. S2 and later prompts should describe the next stage of the same continuous action.
Do not repeat the same action as a new command in every section, because the model may replay it repeatedly.
A complete Russian/English FL2VA prompt guide is included.
REF2VA references
REF2VA supports:
up to 4 image references;
up to 3 video references;
soundtracks from reference videos;
up to 3 standalone audio references.
Enable inputs consecutively inside each type: input 1 before input 2, and input 2 before input 3.
Reference videos are loaded at 24 FPS and limited to the first 360 frames / 15 seconds.
For the first test, use:
one reference;
0.31 MP resolution;
22F context;
0F audio history;
reference_policy = S1 ONLY.
S1 ONLY — recommended
Real image, video, and audio references are presented to the model only during S1. S2–S12 continue through the protected AV latent.
This reduces the risk of later sections suddenly reproducing:
the original reference image;
the original pose or composition;
the reference background;
the original video movement;
the reference audio as a restarted sound.
Use <Picture N>, <Video N>, and <Audio N> tags only in the S1 prompt. Do not use reference tags in S2–S12.
EVERY SECTION
This mode presents all references again during every section.
It can provide stronger reference adherence, but it may also reproduce the original pose, scene, motion, background, or audio as a separate insert. It also requires more processing time and memory.
Each active REF2VA section requires a separate prompt with six blocks:
subject_definitionssummaryretention_analysisdetailed_descriptionoverall_soundscapenon_diegetic_music
A complete Russian/English REF2VA prompt guide with S1 and continuation templates is included.
Recommended starting settings
Resolution: 0.31 MP
Section length: 141F
Context: 22F balanced
Audio history: 0F off
Seed strategy: increment per section
Reference image size: MATCH
REF2VA reference policy: S1 ONLY
Video shift: 12
Audio shift: 3
Increase the resolution and number of references only after a successful test.
Sampling and LoRAs
Available sampling profiles:
Turbo 4 steps
Turbo 8 steps
Native 20 steps
Native 25 steps
The sampling profile must match the selected LoRA. For example, use a 4-step Turbo LoRA with the 4-step profile and an 8-step Turbo LoRA with the 8-step profile.
FL2VA and REF2VA require LoRAs intended for the corresponding model type.
Attention backends
The workflow uses exactly one selected attention backend:
COMFY KITCHEN
SAGE
PYTORCH
Comfy Kitchen is selected in the supplied workflows. If it is unavailable in your ComfyUI build, select SAGE or PYTORCH.
SAGE requires ComfyUI-KJNodes.
Required models
MiniMax H3 models:
Compatible Heretic Text Encoder — choose one:
Optional LoRA sources:
Required custom nodes
ComfyUI-KJNodes — required only for SAGE Attention.
ComfyUI-VideoHelperSuite — required for REF2VA video references.
h3_direct_latent_head — included with this release and required by both workflows.
No latent-upscaler node is required.
Model Preview Override and stability
Model Preview Override is included and enabled so intermediate sampling previews can be monitored during generation.
On some Windows and PyTorch configurations, preview decoding may occasionally cause:
Fatal Python error: Aborted
If this happens:
Refresh the ComfyUI browser page.
Bypass the Model Preview Override node.
Queue the workflow again.
Bypassing Model Preview Override disables only the intermediate preview. It does not change the final generation, prompts, video continuity, audio timeline, or output quality.
Comfy model compiler graph breaks is diagnostic information and does not mean that generation has failed.
The node pack and workflow contracts were tested against ComfyUI 0.34.0.
Optional final post-processing
If additional upscaling or frame interpolation is required, process the completed video afterward with:
It supports RTX Video Super Resolution, DLSS Neural Rendering, and optional frame generation. It is a separate application and is not required for these ComfyUI workflows.
Download DLSS 5 Visual Enhancer
MiniMax H3 Long Video Workflows — V3.0
V3.0 includes two six-section MiniMax H3 long-video workflows and the h3_direct_latent_head custom node.
FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.
REF2VA — supports up to four optional image references across S1–S6 and HQ refinement.
Main features
Direct transfer of video and audio latents between sections.
Exact phase-aligned 22-frame protected AV context.
Accurate 24 FPS video / 40 Hz audio timeline.
Optional audio history: 0F, 24F, 48F, or 72F. Recommended: 48F / 2 seconds.
Final audio is assembled cumulatively, decoded once, and trimmed to the exact video duration.
Automatic section-length and final-duration calculation.
H3 denoise-mask compatibility fix based on ComfyUI PR #15988.
MASTER node for resolution, aspect ratio, quality, continuity, section length, audio history, and S1–S6 prompts.
Individual section checkpoints and final assembled video.
Separate FL2VA and REF2VA prompt-writing guides.
Quality modes
NATIVE — generation without latent upscaling or HQ refinement.
DIRECT — learned latent upscale without diffusion redraw. Recommended for maximum stability.
HQ REFINE — learned latent upscale followed by three-step high-resolution refinement. This can improve detail but may slightly redraw some frames.
Upscaling and HQ refinement do not regenerate the audio.
Turbo LoRA affects S1–S6 and HQ refinement. The secondary motion/style LoRA affects only S1–S6.
Video length
Valid section lengths follow this grid:
section_frames = 5 + 17 × N
Minimum value: 39.
For six sections:
final_frames = 6 × section_frames − 110
The first section keeps its full length. Each following section contributes section_frames − 22 new frames.
Models
MiniMax H3 models:
Heretic Text Encoder — choose one:
LoRA sources:
Learned latent upscaler:
Required custom nodes
h3_direct_latent_head — included in the V3.0 archive.
The external latent-upscaler node must be installed for the workflows to load, even when using NATIVE mode.
Installation
Stop ComfyUI.
Delete the previous
ComfyUI/custom_nodes/h3_direct_latent_headfolder.Copy
h3_direct_latent_headfrom the V3.0 archive intoComfyUI/custom_nodes/.Install the other required custom nodes and models.
Restart ComfyUI and open the FL2VA or REF2VA workflow.
Do not merge files from different versions of the custom node.
Model paths stored in the workflows are examples. Select the corresponding files installed on your system if the paths or filenames differ.
Usage
Configure the workflow through the MASTER node:
quality mode and resolution;
aspect ratio;
continuity profile;
section_frames;audio-history length;
prompts for S1–S6.
FL2VA modes:
T2V: keep START FRAME and END FRAME bypassed.
I2V: enable START FRAME only.
First/last-frame generation: enable both START FRAME and END FRAME.
In REF2VA, all four references are bypassed by default. Enable them consecutively: REF 1, REF 2, REF 3, and REF 4. Use only the corresponding <Picture 1>–<Picture 4> tags in the prompts.
For the first REF2VA test, use REF 1 with the DIRECT quality mode.
Recommended final post-processing
After assembling the final video, I recommend processing it with DLSS 5 Visual Enhancer.
It can improve perceived detail, upscale the completed video, and optionally increase the frame rate using DLSS Frame Generation.
DLSS 5 Visual Enhancer is optional and runs separately from ComfyUI.
Download DLSS 5 Visual Enhancer
My Telegram channel:
MiniMax H3 Long Video Workflows — V2.0

V2.0 includes two six-section long-video workflows:
FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.
REF2VA — supports up to four optional image references connected across all six sections and HQ refinement.
Main features
Direct transfer of video and audio latents between sections.
22-frame AV context for seamless continuation.
Accurate cumulative audio timeline.
Automatic section length and final video duration calculation.
MASTER node for aspect ratio, resolution, quality preset, section length, and S1–S6 prompts.
Native generation or optional learned latent upscale with 3-step HQ refinement.
Individual section checkpoints and final assembled video.
Compatible with Sage Attention, Fused Modulation, and FFN chunking.
Turbo LoRA affects S1–S6 and HQ refinement; the secondary motion/style LoRA affects only S1–S6.
Separate FL2VA and REF2VA prompt-writing guides are included.
Models
MiniMax H3 models:
Heretic Text Encoder — choose one:
Turbo LoRA options:
Optional latent upscaler:
Required custom nodes
h3_direct_latent_head — included in the V2.0 archive.
Copy h3_direct_latent_head into:
ComfyUI/custom_nodes/
Then restart ComfyUI and open the required FL2VA or REF2VA JSON workflow.
Model and LoRA paths stored in the workflow are examples. Select the corresponding files installed on your system if your folder structure or filenames differ.
Usage notes
Edit the MASTER node to select the quality preset, aspect ratio, section_frames, and prompts for S1–S6.
In REF2VA, all four references are bypassed by default. Enable references consecutively starting from REF 1. The included example prompts assume that REF 1 is active.
For REF2VA prompts:
S1: define references and subjects, identity retention, action, camera, and sound.
S2–S6: repeat subject definitions and retention rules, then continue from the transferred latent context without resetting the pose or scene.
See the included FL2VA and REF2VA prompt guides for complete templates and examples.
My Telegram channel:
https://t.me/aibobpublic
V1.0
All Minimax H3 models download from:
https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main
Heretic Text Encoder int8:
https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot/tree/main
Heretic Text Encoder nvfp4:
https://huggingface.co/sakamakismile/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4/tree/main
Turbo loras:
1) https://huggingface.co/DarkRomeo88/MiniMax-H3-turbo-lora-comfyui/tree/main (without larryvrh nodes) and https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main (with larryvrh nodes).
2) Also recommend turbo loras:
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras
3) And:
Required custom nodes:
KJNodes:
https://github.com/kijai/ComfyUI-KJNodes
ComfyUI-sol-attn:
https://github.com/Saganaki22/ComfyUI-sol-attn
I also developed a custom node for transferring encoded latents between sections. This allows the encoded video and audio latents from the previous section to be passed directly to the next section, helping to avoid overcooking during continuation.
For the workflow to work, download the custom node archive and place the folder from the archive into the ComfyUI/custom_nodes/ folder.
You can download it from the drop-down list on this page. There is a separate archive for the workflows and a separate archive for the custom node.
That is, REF2VA formula for all long-video
S1:
REFERENCE BINDING + IDENTITY LOCK + STYLE + ACTION + CAMERA
S2–S6:
CONTINUE FROM LATENT CONTEXT + IDENTITY LOCK FROM OPENING FRAMES + ENVIRONMENT CONTINUITY + ACTION + CAMERA
My TG Channel:
Description
MiniMax H3 Long Video Workflows — V4.0
V4.0 is a complete redesign of the MiniMax H3 long-video workflows. It is smaller, easier to configure, and supports from 1 to 12 sequential sections.
The workflow transfers protected video and audio latent context between sections, allowing MiniMax H3 to generate a longer continuous video instead of unrelated clips.
Included workflows
FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.
REF2VA — supports image, video, video-soundtrack, and standalone audio references.
This release contains two archives:
Workflows + Instruction Prompts — both V4.0 workflows and bilingual Russian/English prompt guides.
h3_direct_latent_head — the shared custom node pack required by both workflows.
Download and install both files.
What changed since V3.0
The large fixed graph was replaced with compact dynamic generation.
The number of sections can be selected from 1 to 12.
You mainly configure only the MASTER and H3 MODEL STACK nodes.
The learned latent upscaler and HQ refinement were removed.
Every section is generated directly at the selected native resolution.
Only one selected model branch and one attention backend are used.
Turbo 4-step, Turbo 8-step, Native 20-step, and Native 25-step profiles are available.
Seed increment per section is enabled by default to reduce repetitive motion.
FL2VA and REF2VA use separate validated prompt formats.
REF2VA now supports image, video, video-audio, and standalone audio references.
REF2VA includes a safer reference policy for long-video continuation.
Final audio is decoded once from the complete cumulative AV timeline.
Model Preview Override is included and enabled.
V4.0 prioritizes native generation, continuity, and a simpler workflow over experimental upscaling and refinement.
How the long-video system works
Each section is a generation window. A protected part of the previous section’s video and audio latent is transferred into the next section.
The recommended overlap is 22 frames. It helps preserve:
the current pose and movement;
character appearance;
camera position and direction;
lighting and scene geometry;
continuous sound and ambience.
With the default settings:
6 sections;
141 frames per section;
22 protected context frames;
24 FPS;
the final video contains 736 frames, or approximately 30.67 seconds.
The first section contributes all 141 frames. Each following section contributes 119 new frames because 22 frames are reused as continuity context.
Installation
Stop ComfyUI completely.
Delete the previous folder:
ComfyUI/custom_nodes/h3_direct_latent_headExtract the new
h3_direct_latent_headfolder into:ComfyUI/custom_nodes/Do not merge files from different node versions.
Install the required external custom nodes.
Place the MiniMax H3 models, text encoder, video VAE, audio VAE, and optional LoRAs in the standard ComfyUI model folders.
Restart ComfyUI.
Open the required FL2VA or REF2VA V4.0 workflow.
H3 MODEL STACK uses portable AUTO selections by default. If several matching files are found, select the correct model or LoRA manually.
FL2VA modes
Control the generation mode using the START and END image inputs:
T2V: START and END are bypassed.
I2V: enable START only.
Last-frame guidance: enable END only.
FL2V: enable both START and END.
Each active section requires a separate prompt with three blocks:
integrated_multimodal_descriptionoverall_soundscapenon_diegetic_music
S1 establishes the subject, scene, camera, movement, and sound. S2 and later prompts should describe the next stage of the same continuous action.
Do not repeat the same action as a new command in every section, because the model may replay it repeatedly.
A complete Russian/English FL2VA prompt guide is included.
REF2VA references
REF2VA supports:
up to 4 image references;
up to 3 video references;
soundtracks from reference videos;
up to 3 standalone audio references.
Enable inputs consecutively inside each type: input 1 before input 2, and input 2 before input 3.
Reference videos are loaded at 24 FPS and limited to the first 360 frames / 15 seconds.
For the first test, use:
one reference;
0.31 MP resolution;
22F context;
0F audio history;
reference_policy = S1 ONLY.
S1 ONLY — recommended
Real image, video, and audio references are presented to the model only during S1. S2–S12 continue through the protected AV latent.
This reduces the risk of later sections suddenly reproducing:
the original reference image;
the original pose or composition;
the reference background;
the original video movement;
the reference audio as a restarted sound.
Use <Picture N>, <Video N>, and <Audio N> tags only in the S1 prompt. Do not use reference tags in S2–S12.
EVERY SECTION
This mode presents all references again during every section.
It can provide stronger reference adherence, but it may also reproduce the original pose, scene, motion, background, or audio as a separate insert. It also requires more processing time and memory.
Each active REF2VA section requires a separate prompt with six blocks:
subject_definitionssummaryretention_analysisdetailed_descriptionoverall_soundscapenon_diegetic_music
A complete Russian/English REF2VA prompt guide with S1 and continuation templates is included.
Recommended starting settings
Resolution: 0.31 MP
Section length: 141F
Context: 22F balanced
Audio history: 0F off
Seed strategy: increment per section
Reference image size: MATCH
REF2VA reference policy: S1 ONLY
Video shift: 12
Audio shift: 3
Increase the resolution and number of references only after a successful test.
Sampling and LoRAs
Available sampling profiles:
Turbo 4 steps
Turbo 8 steps
Native 20 steps
Native 25 steps
The sampling profile must match the selected LoRA. For example, use a 4-step Turbo LoRA with the 4-step profile and an 8-step Turbo LoRA with the 8-step profile.
FL2VA and REF2VA require LoRAs intended for the corresponding model type.
Attention backends
The workflow uses exactly one selected attention backend:
COMFY KITCHEN
SAGE
PYTORCH
Comfy Kitchen is selected in the supplied workflows. If it is unavailable in your ComfyUI build, select SAGE or PYTORCH.
SAGE requires ComfyUI-KJNodes.
Required models
MiniMax H3 models:
Compatible Heretic Text Encoder — choose one:
Optional LoRA sources:
Required custom nodes
ComfyUI-KJNodes — required only for SAGE Attention.
ComfyUI-VideoHelperSuite — required for REF2VA video references.
h3_direct_latent_head — included with this release and required by both workflows.
No latent-upscaler node is required.
Model Preview Override and stability
Model Preview Override is included and enabled so intermediate sampling previews can be monitored during generation.
On some Windows and PyTorch configurations, preview decoding may occasionally cause:
Fatal Python error: Aborted
If this happens:
Refresh the ComfyUI browser page.
Bypass the Model Preview Override node.
Queue the workflow again.
Bypassing Model Preview Override disables only the intermediate preview. It does not change the final generation, prompts, video continuity, audio timeline, or output quality.
Comfy model compiler graph breaks is diagnostic information and does not mean that generation has failed.
The node pack and workflow contracts were tested against ComfyUI 0.34.0.
Optional final post-processing
If additional upscaling or frame interpolation is required, process the completed video afterward with:
It supports RTX Video Super Resolution, DLSS Neural Rendering, and optional frame generation. It is a separate application and is not required for these ComfyUI workflows.

