MiniMax H3 general video generation workflow (VDN 12-step HD edition). Generate high-definition video with text-native audio from keyframes and reference images, without any external audio input.
Main features:
- Inputs: first frame, last frame, up to 3 reference images, plus an optional source video used for frame reference only (its audio is not used)
- Official Duration Planner for text-native audio (17n+5 frame grid)
- OpenVDN DMD8 two-pass pipeline: 8NFE base pass, learned latent x2 upscale to 1920x1088 / 1088x1920 (16:9 or 9:16), 4NFE refine (12-step total)
- Audio audit, AV decode, output trim and safe AV save at approx 1080P / 24fps
Suggested workflow:
1. Provide first/last frames and reference images with the same aspect ratio
2. Optionally load a source video as visual reference only
3. Run the graph; the planner, both passes, upscale and audio audit chain runs automatically, then export the approx 1080P save
Dependencies:
- minimax_h3_fl_ref_hybrid_scheme_a_25-49_bf16.safetensors (H3 ref UNET)
- qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (CLIP)
- minimax_h3_video_vae_fp16.safetensors and minimax_h3_audio_vae_fp32.safetensors
- minimax_h3_latent_upscaler_3d_fp16.safetensors (learned latent x2)
- OpenVDN/vdn-minimax-h3 (DMD8 stage), KSK T8 custom nodes
RunningHub workflow page: https://www.runninghub.ai/zh-cn/post/2108580251193323522?inviteCode=rh-v1111
Bilibili tutorial video: https://www.bilibili.com/video/BV1pTph6EEUv/
RunningHub registration with invite code rh-v1111 (1000 RH coins): https://www.runninghub.ai/?inviteCode=rh-v1111
Description
MiniMax H3 general video generation workflow (VDN 12-step HD edition). Generate high-definition video with text-native audio from keyframes and reference images, without any external audio input.
Main features:
- Inputs: first frame, last frame, up to 3 reference images, plus an optional source video used for frame reference only (its audio is not used)
- Official Duration Planner for text-native audio (17n+5 frame grid)
- OpenVDN DMD8 two-pass pipeline: 8NFE base pass, learned latent x2 upscale to 1920x1088 / 1088x1920 (16:9 or 9:16), 4NFE refine (12-step total)
- Audio audit, AV decode, output trim and safe AV save at approx 1080P / 24fps
Suggested workflow:
1. Provide first/last frames and reference images with the same aspect ratio
2. Optionally load a source video as visual reference only
3. Run the graph; the planner, both passes, upscale and audio audit chain runs automatically, then export the approx 1080P save
Dependencies:
- minimax_h3_fl_ref_hybrid_scheme_a_25-49_bf16.safetensors (H3 ref UNET)
- qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (CLIP)
- minimax_h3_video_vae_fp16.safetensors and minimax_h3_audio_vae_fp32.safetensors
- minimax_h3_latent_upscaler_3d_fp16.safetensors (learned latent x2)
- OpenVDN/vdn-minimax-h3 (DMD8 stage), KSK T8 custom nodes
RunningHub workflow page: https://www.runninghub.ai/zh-cn/post/2108580251193323522?inviteCode=rh-v1111
Bilibili tutorial video: https://www.bilibili.com/video/BV1pTph6EEUv/
RunningHub registration with invite code rh-v1111 (1000 RH coins): https://www.runninghub.ai/?inviteCode=rh-v1111
