CivArchive
    MiniMax H3 Ref2VA Character Replace, OpenVDN DMD8 - v1.0
    NSFW
    Preview 144346587
    Replace the person in an existing video with a new character defined by one reference image - while the original scene, camera movement, lighting and the original audio track stay untouched.

    This workflow is the single-video depth composite design for MiniMax H3. Video 1 and Picture 1 are the only two inputs: no second video, no first/last frame pair.

    How the chain works
    - Video 1 is loaded at 24 fps (124-frame cap in the shipped config) and downscaled to ~0.94 MP on 32-pixel steps (sample: 1152x864, 124 frames).
    - You mark positive points on the target person and negative points on exclusions in the PointsEditor. SeC (SeCVideoSegmentation) tracks the person across the whole video; GrowMask (+8 px) expands the person region.
    - Video Depth Anything (ViT-S) extracts full-frame grayscale depth for every frame.
    - ImageCompositeMasked keeps the original RGB everywhere except the tracked person, which is replaced by its depth rendering. The result is ONE hybrid Video 1: RGB scene + Depth person.
    - MiniMaxH3ReferenceToVideo (Ref2VA) receives Picture 1 as the identity reference and the hybrid video as the structural reference. The built-in subject prompt defines three subjects: the replacement character from Picture 1, the RGB regions to preserve, and the depth region that only guides pose, silhouette, 3D volume, limb ordering and motion.
    - The original audio is encoded with the H3 audio VAE and locked (SolidMask + SetLatentNoiseMask) so the soundtrack passes through without regeneration.
    - Sampling: the MiniMax H3 Hybrid INT8 base is composed with OpenVDN (vdn-minimax-h3, stage_dmd_8nfe), which runs the fixed 8-NFE execution plan in SamplerCustomAdvanced (fixed seed 9527), then video VAE decode and CreateVideo mux in the original audio before SaveVideo.

    Main features:
    - One video + one image input; single-video design, no second clip required
    - Original scene preservation: all unmasked RGB content (background, framing, camera movement, lighting, other people, props) is kept
    - Original audio passthrough: the audio latent is locked and never re-synthesized
    - Fast generation: OpenVDN DMD8 distillation stage, fixed 8 NFE
    - Structural control: the depth region drives body position, pose, limb depth ordering, gesture trajectory and spatial placement; it is never rendered as gray texture
    - No LoRA needed - character identity comes entirely from the Picture 1 reference

    Suggested workflow:
    1. Load only Video 1 (the clip containing the person to replace) and Picture 1 (the new character reference).
    2. In the PointsEditor, mark positive points on the target person and negative points on things that must not be included; let SeC track the whole video.
    3. Inspect the tracked mask before running. GrowMask defaults to 8 - adjust it so the background around the person is not converted to depth.
    4. Queue the prompt. The output MP4 keeps the scene and the original audio while the person is replaced with the Picture 1 identity.

    Model and node dependencies:
    - UNet: minimax_h3_hybrid_fl2va_ref2va_zs05_b25-49_int8.safetensors (MiniMax H3 Hybrid INT8, OpenVDN structural base)
    - CLIP: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (Qwen3VL 32B NVFP4, type: minimax)
    - Video VAE: minimax_h3_video_vae_int8_convrot.safetensors
    - Audio VAE: minimax_h3_audio_vae_fp32.safetensors
    - Depth: video_depth_anything_vits.pth (Video Depth Anything ViT-S)
    - SeC model loaded via SeCModelLoader (bfloat16 / auto device)
    - Custom nodes: minimax-h3-audio-T8 (MiniMaxH3ReferenceToVideo, MiniMaxH3VDNModelComposerT8Advanced, MiniMaxH3VDNExecutionPlanT8Advanced, LTXVAudioVAEEncode, LTXVConcatAVLatent, LTXVSeparateAVLatent), ComfyUI-VideoHelperSuite (VHS_LoadVideo), ComfyUI-VideoDepthAnything (LoadVideoDepthAnythingModel, VideoDepthAnythingProcess), SeC nodes (SeCModelLoader, SeCVideoSegmentation), KJNodes (PointsEditor)

    Inputs / outputs:
    - Inputs: 1 video (MP4, up to 124 frames in the shipped config) + 1 character reference image
    - Output: MP4 at ~0.94 MP (sample 1152x864) at the source frame rate (24 fps), original audio track preserved

    Limitations:
    - Designed for one main person per video; multiple people require separate masked passes
    - The shipped config caps at 124 frames - raise frame_load_cap for longer clips (VRAM dependent)
    - Fixed 8 NFE (DMD8 stage): fast, distilled quality
    - The grayscale depth area is control information only - the prompt forbids it from appearing as visible gray

    Bypassed / unused nodes: none - all 31 nodes are active in this workflow.

    Links (to be added as they go live - none provided for this batch, none fabricated):
    - YouTube: pending
    - Bilibili: pending
    - RunningHub: pending

    Description

    Replace the person in an existing video with a new character defined by one reference image - while the original scene, camera movement, lighting and the original audio track stay untouched.

    This workflow is the single-video depth composite design for MiniMax H3. Video 1 and Picture 1 are the only two inputs: no second video, no first/last frame pair.

    How the chain works
    - Video 1 is loaded at 24 fps (124-frame cap in the shipped config) and downscaled to ~0.94 MP on 32-pixel steps (sample: 1152x864, 124 frames).
    - You mark positive points on the target person and negative points on exclusions in the PointsEditor. SeC (SeCVideoSegmentation) tracks the person across the whole video; GrowMask (+8 px) expands the person region.
    - Video Depth Anything (ViT-S) extracts full-frame grayscale depth for every frame.
    - ImageCompositeMasked keeps the original RGB everywhere except the tracked person, which is replaced by its depth rendering. The result is ONE hybrid Video 1: RGB scene + Depth person.
    - MiniMaxH3ReferenceToVideo (Ref2VA) receives Picture 1 as the identity reference and the hybrid video as the structural reference. The built-in subject prompt defines three subjects: the replacement character from Picture 1, the RGB regions to preserve, and the depth region that only guides pose, silhouette, 3D volume, limb ordering and motion.
    - The original audio is encoded with the H3 audio VAE and locked (SolidMask + SetLatentNoiseMask) so the soundtrack passes through without regeneration.
    - Sampling: the MiniMax H3 Hybrid INT8 base is composed with OpenVDN (vdn-minimax-h3, stage_dmd_8nfe), which runs the fixed 8-NFE execution plan in SamplerCustomAdvanced (fixed seed 9527), then video VAE decode and CreateVideo mux in the original audio before SaveVideo.

    Main features:
    - One video + one image input; single-video design, no second clip required
    - Original scene preservation: all unmasked RGB content (background, framing, camera movement, lighting, other people, props) is kept
    - Original audio passthrough: the audio latent is locked and never re-synthesized
    - Fast generation: OpenVDN DMD8 distillation stage, fixed 8 NFE
    - Structural control: the depth region drives body position, pose, limb depth ordering, gesture trajectory and spatial placement; it is never rendered as gray texture
    - No LoRA needed - character identity comes entirely from the Picture 1 reference

    Suggested workflow:
    1. Load only Video 1 (the clip containing the person to replace) and Picture 1 (the new character reference).
    2. In the PointsEditor, mark positive points on the target person and negative points on things that must not be included; let SeC track the whole video.
    3. Inspect the tracked mask before running. GrowMask defaults to 8 - adjust it so the background around the person is not converted to depth.
    4. Queue the prompt. The output MP4 keeps the scene and the original audio while the person is replaced with the Picture 1 identity.

    Model and node dependencies:
    - UNet: minimax_h3_hybrid_fl2va_ref2va_zs05_b25-49_int8.safetensors (MiniMax H3 Hybrid INT8, OpenVDN structural base)
    - CLIP: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (Qwen3VL 32B NVFP4, type: minimax)
    - Video VAE: minimax_h3_video_vae_int8_convrot.safetensors
    - Audio VAE: minimax_h3_audio_vae_fp32.safetensors
    - Depth: video_depth_anything_vits.pth (Video Depth Anything ViT-S)
    - SeC model loaded via SeCModelLoader (bfloat16 / auto device)
    - Custom nodes: minimax-h3-audio-T8 (MiniMaxH3ReferenceToVideo, MiniMaxH3VDNModelComposerT8Advanced, MiniMaxH3VDNExecutionPlanT8Advanced, LTXVAudioVAEEncode, LTXVConcatAVLatent, LTXVSeparateAVLatent), ComfyUI-VideoHelperSuite (VHS_LoadVideo), ComfyUI-VideoDepthAnything (LoadVideoDepthAnythingModel, VideoDepthAnythingProcess), SeC nodes (SeCModelLoader, SeCVideoSegmentation), KJNodes (PointsEditor)

    Inputs / outputs:
    - Inputs: 1 video (MP4, up to 124 frames in the shipped config) + 1 character reference image
    - Output: MP4 at ~0.94 MP (sample 1152x864) at the source frame rate (24 fps), original audio track preserved

    Limitations:
    - Designed for one main person per video; multiple people require separate masked passes
    - The shipped config caps at 124 frames - raise frame_load_cap for longer clips (VRAM dependent)
    - Fixed 8 NFE (DMD8 stage): fast, distilled quality
    - The grayscale depth area is control information only - the prompt forbids it from appearing as visible gray

    Bypassed / unused nodes: none - all 31 nodes are active in this workflow.

    Links (to be added as they go live - none provided for this batch, none fabricated):
    - YouTube: pending
    - Bilibili: pending
    - RunningHub: pending
    Workflows
    Other

    Details

    Downloads
    81
    Platform
    CivitAI
    Platform Status
    Available
    Created
    10/1/2026
    Updated
    10/2/2026
    Deleted
    -

    Files

    minimaxH3Ref2vaCharacter_v10.json

    Mirrors