This workflow is a MiniMax H3 audio-driven digital human setup for a single character. It is designed for turning one character reference and one audio track into a direct speaking or performing avatar video, with the audio driving the timing while the prompt defines acting style, camera language, and scene mood.
The workflow uses the MiniMax H3 local ComfyUI graph with the connected model, VAE, encoder, sampler, resolution, and final video output chain shown in the JSON.
Main features:
- MiniMax H3 reference-to-video route for a single speaking subject
- Connected MiniMax H3 Ref2VA diffusion model with Qwen3-VL text encoding
- Video VAE and audio VAE routes are both connected to the generation chain
- Audio input drives speech timing, performance rhythm, and mouth movement
- 16:9 widescreen output at the connected 0.9 megapixel resolution route
- 20-step simple scheduler with res_multistep sampling
- Final video combine route keeps the generated performance and audio together
Suggested workflow:
Use one clean character image or first frame and a clearly recorded audio file. Keep the prompt focused on one subject, visible face, readable lips, emotional tone, and camera framing. For better stability, avoid asking for extra people, large identity changes, or fast scene cuts during the first test.
RunningHub Workflow
Try the workflow online right now - no installation required.
Workflow: https://www.runninghub.ai/post/2085293371178201089?inviteCode=rh-v1111
If the results meet your expectations, you can later deploy it locally for customization.
Fan Benefits: Register to get 1000 points + daily login 100 points - enjoy 4090 performance and 48 GB super power!
Bilibili Updates (Mainland China & Asia-Pacific)
If you're in the Asia-Pacific region, you can watch the video below to see the workflow demonstration and creative breakdown.
Bilibili Video: https://www.bilibili.com/video/BV1kUuW6FEVs/
Support Me on Ko-fi
If you find my content helpful and want to support future creations, you can buy me a coffee.
Every bit of support helps me keep creating.
Ko-fi: https://ko-fi.com/aiksk
Business Contact
For collaboration or inquiries, please contact aiksk95 on WeChat.
打开下方链接即可在线体验,无需安装。
工作流:https://www.runninghub.ai/post/2085293371178201089?inviteCode=rh-v1111
如果你觉得效果理想,也可以在本地进行自定义部署。
粉丝福利:注册即可领取 1000 积分,每日登录再领 100 积分,体验 4090 和 48 GB 大显存性能。
B站视频(中国大陆及亚太地区)
如果你在中国大陆或亚太地区,可以通过下面的视频查看工作流的实测效果与创作思路。
B站视频:https://www.bilibili.com/video/BV1kUuW6FEVs/
我会在夸克网盘持续更新模型资源:
https://pan.quark.cn/s/07bdc81784ce
Description
This workflow is a MiniMax H3 audio-driven digital human setup for a single character. It is designed for turning one character reference and one audio track into a direct speaking or performing avatar video, with the audio driving the timing while the prompt defines acting style, camera language, and scene mood.
The workflow uses the MiniMax H3 local ComfyUI graph with the connected model, VAE, encoder, sampler, resolution, and final video output chain shown in the JSON.
Main features:
- MiniMax H3 reference-to-video route for a single speaking subject
- Connected MiniMax H3 Ref2VA diffusion model with Qwen3-VL text encoding
- Video VAE and audio VAE routes are both connected to the generation chain
- Audio input drives speech timing, performance rhythm, and mouth movement
- 16:9 widescreen output at the connected 0.9 megapixel resolution route
- 20-step simple scheduler with res_multistep sampling
- Final video combine route keeps the generated performance and audio together
Suggested workflow:
Use one clean character image or first frame and a clearly recorded audio file. Keep the prompt focused on one subject, visible face, readable lips, emotional tone, and camera framing. For better stability, avoid asking for extra people, large identity changes, or fast scene cuts during the first test.
RunningHub Workflow
Try the workflow online right now - no installation required.
Workflow: https://www.runninghub.ai/post/2085293371178201089?inviteCode=rh-v1111
If the results meet your expectations, you can later deploy it locally for customization.
Fan Benefits: Register to get 1000 points + daily login 100 points - enjoy 4090 performance and 48 GB super power!
Bilibili Updates (Mainland China & Asia-Pacific)
If you're in the Asia-Pacific region, you can watch the video below to see the workflow demonstration and creative breakdown.
Bilibili Video: https://www.bilibili.com/video/BV1kUuW6FEVs/
Support Me on Ko-fi
If you find my content helpful and want to support future creations, you can buy me a coffee.
Every bit of support helps me keep creating.
Ko-fi: https://ko-fi.com/aiksk
Business Contact
For collaboration or inquiries, please contact aiksk95 on WeChat.
打开下方链接即可在线体验,无需安装。
工作流:https://www.runninghub.ai/post/2085293371178201089?inviteCode=rh-v1111
如果你觉得效果理想,也可以在本地进行自定义部署。
粉丝福利:注册即可领取 1000 积分,每日登录再领 100 积分,体验 4090 和 48 GB 大显存性能。
B站视频(中国大陆及亚太地区)
如果你在中国大陆或亚太地区,可以通过下面的视频查看工作流的实测效果与创作思路。
B站视频:https://www.bilibili.com/video/BV1kUuW6FEVs/
我会在夸克网盘持续更新模型资源:
https://pan.quark.cn/s/07bdc81784ce
FAQ
minimax h3
workflows
ai video
runninghub
workflow
comfyui
lip sync
talking avatar
audio driven video
single avatar
digital human
Details
Downloads
108
Platform
CivitAI
Platform Status
Available
Created
8/7/2026
Updated
8/12/2026
Deleted
-
