One workflow for everything MiniMax H3 can generate.
Full walkthrough on YouTube:
https://youtu.be/yUQ3oa3jULU
H3 makes video and audio at the same time, and it ships as two different UNets
with two different conditioning nodes. This graph has all seven modes in it, so
you don't need a separate workflow per job:
01 Text to Video prompt only
02 Image to Video one image as the first frame
03 First / Last Frame image A to image B
04 Reference to Video up to 5 images + 3 videos + 3 audio clips
05 Motion Tracking DWPose skeleton
06 Motion Tracking DensePose body-surface map
07 Motion Tracking SCAIL / NLF 3D mesh
Lip sync is separate from the mode list and works on top of any of the seven.
To use it: pick a UNet in the (1) Model panel, pick a mode in the (2) panel,
hit run. Everything else is optional. The notes inside the graph cover the
model downloads, the resolution table, prompt format and the speed settings.
Full walkthrough on YouTube:
<<ここに YouTube のリンク>>
Two files are attached. The graphs are identical - only the language of the
notes and group labels differs. Take whichever one you read faster.
*_en_v1.json English
*_ja_v1.json 日本語
Custom nodes: rgthree-comfy, ComfyUI-KJNodes, ComfyUI-VideoHelperSuite,
WhatDreamsCost-ComfyUI. Modes 05-07 also need the DWPose / DensePose / NLF
preprocessors. Model download links are in the notes inside the graph.
Please don't change the group colours - the panels find what they control by
colour, so a recoloured group drops out of its panel.