I merged the official fl2va workflows.(minimax-h3_fl2va-Mix)
you can process t2va,i2va(first_image),l2va(last_image),fl2va in one workflow now.
Reference output:
RTX5080 laptop 16G vRAM + 64G RAM T2V&I2V(pruned_int8_conv; 8s*1376*768): Prompt executed in 00:20:12
T2V use the official example prompt.
Requirements:
Models:
- minimax_h3_fl2va_pruned_int8_convrot.safetensors, Put in: ComfyUI\models\diffusion_models
- qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors; (CLIP) Put in: ComfyUI\models\text_encoders
- VAE Put in: ComfyUI\models\vae
- minimax_h3_video_vae_fp16.safetensors
https://huggingface.co/Comfy-Org/MiniMax-H3/resolve/main/vae/minimax_h3_video_vae_fp16.safetensors
- minimax_h3_audio_vae_fp32.safetensors
https://huggingface.co/Comfy-Org/MiniMax-H3/resolve/main/vae/minimax_h3_audio_vae_fp32.safetensors
ComfyUI Nodes:
- rgthree-comfy
- ComfyUI-KJNodes
- ComfyUI-VideoHelperSuite
If you have higher performance hardware, you can choose higher quantization models.


