CivArchive
    MiniMax H3 Role Swap Face Refine - v1.0
    Preview 144485617


    This workflow is the flagship role swap / reference editing route in the MiniMax H3 swap episode. It replaces the character in a reference video with the character defined by a reference image, then runs an independent face refine pass over the generated clip to clean up facial detail and identity stability before the final video is exported.

    The inspected graph uses the MiniMax H3 hybrid FL2VA/Ref2VA INT8 base model (minimax_h3_hybrid_fl2va_ref2va_zs05_b25-49_int8.safetensors), Qwen3-VL 32B NVFP4/AWQ text encoding, MiniMax H3 video VAE (INT8 ConvRot) and audio VAE (FP32), and OpenVDN DMD8 execution planning with custom samplers. Depth control comes from Video Depth Anything ViT-S (518x1280 fp16, gray), subject tracking comes from SeC full-clip segmentation with feathered mask compositing, and the swap is driven by the MiniMaxH3ReferenceToVideo control node that combines the Picture 1 identity reference with the single reference video depth composite (768x768, 362-frame setting). Reference audio is encoded through the LTXV audio VAE and the final clip is packaged at a unified 24 FPS with a safe audio channel.

    On top of the base swap, the ADDON Face chain detects faces with YuNet (ONNX face detection), crops a 512 AV latent window around the tracked face, re-samples it with a separate minimax_h3_fl2va_int8_convrot model using a low-denoise 12-step schedule (0.45 denoise, dual_clock_euler / native_flow), and stitches the refined crop back into the full frame. The Stitch + Audit node is a diagnostic output and is not part of the final deliverable.

    This package is meant for ComfyUI and RunningHub users who want a ready-made character replacement graph with a built-in face correction stage, instead of wiring loaders, encoders, VAEs, depth, segmentation, and face refine nodes by hand. It is especially useful when the base swap works but facial detail or identity drift needs an extra pass.

    Main features:

    - Character replacement driven by a multi-view Picture 1 identity reference plus reference video depth composite
    - Video Depth Anything ViT-S whole-frame depth and SeC full-clip segmentation with feathered compositing
    - MiniMax H3 hybrid FL2VA/Ref2VA INT8 base with OpenVDN DMD8 execution planning
    - Independent face refine stage: YuNet detection, 512 AV latent crop, low-denoise 12-step re-sample, stitch back
    - Face tracking preview and refined crop preview outputs for troubleshooting
    - Qwen3-VL 32B NVFP4/AWQ text encoding with MiniMax H3 video and audio VAEs
    - Final video packaged at unified 24 FPS with a safe audio channel; dual-instance graph included

    Suggested workflow:

    Start with the base swap settings from the episode video and keep the same reference image, prompt, duration, and seed when comparing variants. Check the face tracking preview first to confirm the detector locks onto the subject, then review the refined face crop output before judging the final clip. If the face still drifts after the refine pass, improve the reference image quality or the crop window before changing model or sampler settings.

    For fair comparisons across the three variants of this episode, run one short test first with identical inputs, then review identity stability, motion consistency, and background drift before scaling up to full-length renders.

    RunningHub Workflow

    Try the workflow online right now - no installation required.
    Workflow: https://www.runninghub.ai/zh-cn/post/2106050577843560449?inviteCode=rh-v1111

    If the results meet your expectations, you can later deploy it locally for customization.

    Fan Benefits: Register to get 1000 points + daily login 100 points - enjoy 4090 performance and 48 GB super power!

    Bilibili Updates (Mainland China & Asia-Pacific)

    If you're in the Asia-Pacific region, you can watch the video below to see the workflow demonstration and creative breakdown.
    Bilibili Video: https://www.bilibili.com/video/BV1VjaD6jEU3/

    Support Me on Ko-fi

    If you find my content helpful and want to support future creations, you can buy me a coffee.
    Every bit of support helps me keep creating.
    Ko-fi: https://ko-fi.com/aiksk

    Business Contact

    For collaboration or inquiries, please contact aiksk95 on WeChat.

    打开下方链接即可在线体验,无需安装。
    工作流:https://www.runninghub.ai/zh-cn/post/2106050577843560449?inviteCode=rh-v1111

    如果你觉得效果理想,也可以在本地进行自定义部署。

    粉丝福利:注册即可领取 1000 积分,每日登录再领 100 积分,体验 4090 和 48GB 大显存性能。

    B站视频(中国大陆及亚太地区)

    如果你在中国大陆或亚太地区,可以通过下面的视频查看工作流的实测效果与创作思路。
    B站视频:https://www.bilibili.com/video/BV1VjaD6jEU3/

    我会在夸克网盘持续更新模型资源:
    https://pan.quark.cn/s/9ec095b95838

    Description



    This workflow is the flagship role swap / reference editing route in the MiniMax H3 swap episode. It replaces the character in a reference video with the character defined by a reference image, then runs an independent face refine pass over the generated clip to clean up facial detail and identity stability before the final video is exported.

    The inspected graph uses the MiniMax H3 hybrid FL2VA/Ref2VA INT8 base model (minimax_h3_hybrid_fl2va_ref2va_zs05_b25-49_int8.safetensors), Qwen3-VL 32B NVFP4/AWQ text encoding, MiniMax H3 video VAE (INT8 ConvRot) and audio VAE (FP32), and OpenVDN DMD8 execution planning with custom samplers. Depth control comes from Video Depth Anything ViT-S (518x1280 fp16, gray), subject tracking comes from SeC full-clip segmentation with feathered mask compositing, and the swap is driven by the MiniMaxH3ReferenceToVideo control node that combines the Picture 1 identity reference with the single reference video depth composite (768x768, 362-frame setting). Reference audio is encoded through the LTXV audio VAE and the final clip is packaged at a unified 24 FPS with a safe audio channel.

    On top of the base swap, the ADDON Face chain detects faces with YuNet (ONNX face detection), crops a 512 AV latent window around the tracked face, re-samples it with a separate minimax_h3_fl2va_int8_convrot model using a low-denoise 12-step schedule (0.45 denoise, dual_clock_euler / native_flow), and stitches the refined crop back into the full frame. The Stitch + Audit node is a diagnostic output and is not part of the final deliverable.

    This package is meant for ComfyUI and RunningHub users who want a ready-made character replacement graph with a built-in face correction stage, instead of wiring loaders, encoders, VAEs, depth, segmentation, and face refine nodes by hand. It is especially useful when the base swap works but facial detail or identity drift needs an extra pass.

    Main features:

    - Character replacement driven by a multi-view Picture 1 identity reference plus reference video depth composite
    - Video Depth Anything ViT-S whole-frame depth and SeC full-clip segmentation with feathered compositing
    - MiniMax H3 hybrid FL2VA/Ref2VA INT8 base with OpenVDN DMD8 execution planning
    - Independent face refine stage: YuNet detection, 512 AV latent crop, low-denoise 12-step re-sample, stitch back
    - Face tracking preview and refined crop preview outputs for troubleshooting
    - Qwen3-VL 32B NVFP4/AWQ text encoding with MiniMax H3 video and audio VAEs
    - Final video packaged at unified 24 FPS with a safe audio channel; dual-instance graph included

    Suggested workflow:

    Start with the base swap settings from the episode video and keep the same reference image, prompt, duration, and seed when comparing variants. Check the face tracking preview first to confirm the detector locks onto the subject, then review the refined face crop output before judging the final clip. If the face still drifts after the refine pass, improve the reference image quality or the crop window before changing model or sampler settings.

    For fair comparisons across the three variants of this episode, run one short test first with identical inputs, then review identity stability, motion consistency, and background drift before scaling up to full-length renders.

    RunningHub Workflow

    Try the workflow online right now - no installation required.
    Workflow: https://www.runninghub.ai/zh-cn/post/2106050577843560449?inviteCode=rh-v1111

    If the results meet your expectations, you can later deploy it locally for customization.

    Fan Benefits: Register to get 1000 points + daily login 100 points - enjoy 4090 performance and 48 GB super power!

    Bilibili Updates (Mainland China & Asia-Pacific)

    If you're in the Asia-Pacific region, you can watch the video below to see the workflow demonstration and creative breakdown.
    Bilibili Video: https://www.bilibili.com/video/BV1VjaD6jEU3/

    Support Me on Ko-fi

    If you find my content helpful and want to support future creations, you can buy me a coffee.
    Every bit of support helps me keep creating.
    Ko-fi: https://ko-fi.com/aiksk

    Business Contact

    For collaboration or inquiries, please contact aiksk95 on WeChat.

    打开下方链接即可在线体验,无需安装。
    工作流:https://www.runninghub.ai/zh-cn/post/2106050577843560449?inviteCode=rh-v1111

    如果你觉得效果理想,也可以在本地进行自定义部署。

    粉丝福利:注册即可领取 1000 积分,每日登录再领 100 积分,体验 4090 和 48GB 大显存性能。

    B站视频(中国大陆及亚太地区)

    如果你在中国大陆或亚太地区,可以通过下面的视频查看工作流的实测效果与创作思路。
    B站视频:https://www.bilibili.com/video/BV1VjaD6jEU3/

    我会在夸克网盘持续更新模型资源:
    https://pan.quark.cn/s/9ec095b95838
    Workflows
    Other

    Details

    Downloads
    10
    Platform
    CivitAI
    Platform Status
    Available
    Created
    10/2/2026
    Updated
    10/2/2026
    Deleted
    -

    Files

    minimaxH3RoleSwapFace_v10.json

    Mirrors