Masked video inpainting on Wan 2.2 VACE, where the mask comes from a text prompt
instead of manual painting. Say what to replace, say what replaces it, hit run.
What makes it different
Most VACE inpaint workflows need you to paint a mask on frame 0 and propagate it.
This one uses SAM 3.1 to detect and track the target across the whole clip from a
text prompt, so it runs once, start to finish. No painting, no halt, no second pass.
For manual masking use the MatAnyone version instead.
Features
INT8 ConvRot models throughout, native on Ampere and newer, no FP8 emulation
Chunked sampling loop, so clip length isn't bounded by VRAM
Optional reference image to guide appearance (off by default, one toggle)
Optional 4x upscale with per-source-type model recommendations
Runs on 4-step Lightning distills, a few seconds per chunk rather than minutes
Full on-canvas documentation covering every knob, all download links, and
mask-tuning guidance per task type
The rule that matters: VACE can only paint inside the mask. If the new thing is
bigger than what it replaces, expand the mask or it gets squashed into the old
silhouette.
All node packs are in the ComfyUI Manager registry. Model download links and full
documentation are in the Guide - How To Use and Model Links notes on the canvas.
Description
First public release.
Text-prompted masked video inpainting on Wan 2.2 Fun VACE. SAM 3.1 detects and
tracks the target from a plain noun, so the whole clip runs in a single pass —
no mask painting, no second run.
Runs on INT8 ConvRot models and 4-step Lightning distills.
