MiniMax H3 Long Video Workflows — V2.0

V2.0 includes two six-section long-video workflows:
FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.
REF2VA — supports up to four optional image references connected across all six sections and HQ refinement.
Main features
Direct transfer of video and audio latents between sections.
22-frame AV context for seamless continuation.
Accurate cumulative audio timeline.
Automatic section length and final video duration calculation.
MASTER node for aspect ratio, resolution, quality preset, section length, and S1–S6 prompts.
Native generation or optional learned latent upscale with 3-step HQ refinement.
Individual section checkpoints and final assembled video.
Compatible with Sage Attention, Fused Modulation, and FFN chunking.
Turbo LoRA affects S1–S6 and HQ refinement; the secondary motion/style LoRA affects only S1–S6.
Separate FL2VA and REF2VA prompt-writing guides are included.
Models
MiniMax H3 models:
Heretic Text Encoder — choose one:
Turbo LoRA options:
Optional latent upscaler:
Required custom nodes
h3_direct_latent_head — included in the V2.0 archive.
Copy h3_direct_latent_head into:
ComfyUI/custom_nodes/
Then restart ComfyUI and open the required FL2VA or REF2VA JSON workflow.
Model and LoRA paths stored in the workflow are examples. Select the corresponding files installed on your system if your folder structure or filenames differ.
Usage notes
Edit the MASTER node to select the quality preset, aspect ratio, section_frames, and prompts for S1–S6.
In REF2VA, all four references are bypassed by default. Enable references consecutively starting from REF 1. The included example prompts assume that REF 1 is active.
For REF2VA prompts:
S1: define references and subjects, identity retention, action, camera, and sound.
S2–S6: repeat subject definitions and retention rules, then continue from the transferred latent context without resetting the pose or scene.
See the included FL2VA and REF2VA prompt guides for complete templates and examples.
My Telegram channel:
https://t.me/aibobpublic
V1.0
All Minimax H3 models download from:
https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main
Heretic Text Encoder int8:
https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot/tree/main
Heretic Text Encoder nvfp4:
https://huggingface.co/sakamakismile/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4/tree/main
Turbo loras:
1) https://huggingface.co/DarkRomeo88/MiniMax-H3-turbo-lora-comfyui/tree/main (without larryvrh nodes) and https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main (with larryvrh nodes).
2) Also recommend turbo loras:
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras
3) And:
Required custom nodes:
KJNodes:
https://github.com/kijai/ComfyUI-KJNodes
ComfyUI-sol-attn:
https://github.com/Saganaki22/ComfyUI-sol-attn
I also developed a custom node for transferring encoded latents between sections. This allows the encoded video and audio latents from the previous section to be passed directly to the next section, helping to avoid overcooking during continuation.
For the workflow to work, download the custom node archive and place the folder from the archive into the ComfyUI/custom_nodes/ folder.
You can download it from the drop-down list on this page. There is a separate archive for the workflows and a separate archive for the custom node.
That is, REF2VA formula for all long-video
S1:
REFERENCE BINDING + IDENTITY LOCK + STYLE + ACTION + CAMERA
S2–S6:
CONTINUE FROM LATENT CONTEXT + IDENTITY LOCK FROM OPENING FRAMES + ENVIRONMENT CONTINUITY + ACTION + CAMERA
My TG Channel:
Description
MiniMax H3 Long Video Workflows — V2.0
V2.0 includes two six-section long-video workflows:
FL2VA — supports T2V, I2V, last-frame guidance, and first/last-frame generation.
REF2VA — supports up to four optional image references connected across all six sections and HQ refinement.
Main features
Direct transfer of video and audio latents between sections.
22-frame AV context for seamless continuation.
Accurate cumulative audio timeline.
Automatic section length and final video duration calculation.
MASTER node for aspect ratio, resolution, quality preset, section length, and S1–S6 prompts.
Native generation or optional learned latent upscale with 3-step HQ refinement.
Individual section checkpoints and final assembled video.
Compatible with Sage Attention, Fused Modulation, and FFN chunking.
Turbo LoRA affects S1–S6 and HQ refinement; the secondary motion/style LoRA affects only S1–S6.
Separate FL2VA and REF2VA prompt-writing guides are included.
Models
MiniMax H3 models:
https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main
Heretic Text Encoder — choose one:
[INT8 ConvRot https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot/tree/main
NVFP4 https://huggingface.co/sakamakismile/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4/tree/main
Turbo LoRA options:
DarkRomeo88 — ComfyUI version https://huggingface.co/DarkRomeo88/MiniMax-H3-turbo-lora-comfyui/tree/main
LarryVRH version https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main
Kijai experimental LoRAs https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras
LightX2V REF2VA Turbo https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors
Optional latent upscaler:
Upscaler model https://huggingface.co/LBH-123-AI/Minimax_h3_latent_Upscaler/tree/main
ComfyUI upscaler node https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler
Required custom nodes
ComfyUI-KJNodes https://github.com/kijai/ComfyUI-KJNodes
ComfyUI-sol-attn https://github.com/Saganaki22/ComfyUI-sol-attn
h3_direct_latent_head — included in the V2.0 archive.
Copy h3_direct_latent_head into:
ComfyUI/custom_nodes/
Then restart ComfyUI and open the required FL2VA or REF2VA JSON workflow.
Model and LoRA paths stored in the workflow are examples. Select the corresponding files installed on your system if your folder structure or filenames differ.
Usage notes
Edit the MASTER node to select the quality preset, aspect ratio, section_frames, and prompts for S1–S6.
In REF2VA, all four references are bypassed by default. Enable references consecutively starting from REF 1. The included example prompts assume that REF 1 is active.
For REF2VA prompts:
S1: define references and subjects, identity retention, action, camera, and sound.
S2–S6: repeat subject definitions and retention rules, then continue from the transferred latent context without resetting the pose or scene.
See the included FL2VA and REF2VA prompt guides for complete templates and examples.
My Telegram channel:
