CivArchive
    SCAIL-2 GGUF MOTION TRANSFER Reference Image to Video Vid2Vid + MultiGPU - v2.0
    NSFW
    Use Worflow Now for Free: https://www.floyo.ai/workflows/scail-2-bf16-loop-fulllength-motion--in4ss5zsvw3r
    
    SCAIL-2 GGUF MOTION TRANSFER Reference Image to Video + MultiGPU
    
    Turn a single character image into a fully animated video that copies the motion of any driving clip — at the **full length** of your input video, not just a fixed 5-second window. Built on the **Wan 2.1 SCAIL-2** model in quantized **GGUF** format so it runs on consumer GPUs, with optional **dual-GPU weight offloading** to keep speeds up.
    
    ---
    
     ✨ What this workflow does
    
    Feed it **one reference image** (your character) and **one driving video** (the motion). SAM3 automatically tracks and masks the subject, the pose from the driving video drives the animation, and a **real chunked sampling loop** generates the entire clip — stitching sliding windows together with color-matching so there are no harsh seams between chunks.
    
    - **Full-length output** — a sliding-window loop (81-frame initial window + 76-frame continuation windows) covers your whole driving video. No 5-second cap.
    - **Motion transfer** — the subject in your reference image performs the exact motion of the driving clip.
    - **Automatic subject masking** — SAM3.1 tracking isolates the character; no manual rotoscoping.
    - **GGUF quantized model** — Q4_K_M weights fit comfortably in consumer VRAM.
    - **Optional 2nd-GPU offload** — push ~10 GB of model weights to a second GPU instead of slow CPU offload.
    - **Built-in side-by-side comparison output** — see reference vs. result in one render.
    - **Organized & documented** — color-coded node groups and an on-canvas README note with every download link.
    
    ---
    
     🎬 How to use it
    
    1. **Reference image** → load your character in the `LoadImage` node (INPUTS group).
    2. **Driving video** → load your motion clip in `VHS_LoadVideo` (INPUTS group). Leave `custom_width = 480` — it keeps system RAM low and matches the working resolution.
    3. **Prompts** → describe the scene in the positive prompt and what to avoid in the negative (PROMPTS group).
    4. **Press Run.** The output appears **only after the loop finishes** — there are no mid-run previews (this is normal, not a freeze). The OUTPUT group holds the final stitched video; the COMPARISON group shows the side-by-side.
    
    **Speed tip:** set `select_every_nth = 2` on `VHS_LoadVideo` to roughly halve render time at half the temporal resolution. You can also lower the sampler steps.
    
    ---
    
     🖥️ Single GPU vs. Dual GPU (model switcher built in)
    
    The workflow includes two model loaders feeding an **Any Switch (rgthree)** "Model Switcher":
    
    - **GGUF Loader – MULTI GPU (default)** — offloads ~10 GB of weights to your **second GPU** (`cuda:1`), keeping compute on `cuda:0`. Dramatically faster than CPU offload.
    - **GGUF Loader – SINGLE GPU** — standard single-GPU GGUF loading.
    
    **Switching is manual** (ComfyUI can't auto-detect GPU count). Use **Ctrl+B** to bypass the one you don't want — keep **exactly one** active:
    
    - **Two GPUs:** leave as shipped → MultiGPU loader active, single-GPU loader bypassed.
    - **One GPU:** bypass the MultiGPU loader and un-bypass the single-GPU loader. (If you leave the MultiGPU loader active with only one GPU, it will error trying to reach the missing `cuda:1`.)
    
    > Tip: on the MultiGPU loader you can tune `virtual_vram_gb` (default 10) — lower it if your 2nd GPU OOMs, raise it if it has spare room. `donor_device` can also be set to `cpu` for a single-GPU fallback without bypassing.
    
    ---
    
     📦 Required models & paths
    
    Place these under your ComfyUI `models/` folder:
    
    ```
    ComfyUI/
    └── models/
        ├── unet/  (or diffusion_models/)
        │   └── SCAIL-2-Q4_K_M.gguf            ← supply your own GGUF
        ├── text_encoders/
        │   └── umt5_xxl_fp8_e4m3fn_scaled.safetensors
        ├── clip_vision/
        │   └── clip_vision_h.safetensors
        ├── vae/
        │   └── wan_2.1_vae.safetensors
        ├── loras/
        │   └── Wan21_I2V_14B_lightx2v_cfg_step_distill_lora_rank64.safetensors
        └── checkpoints/
            └── sam3.1_multiplex_fp16.safetensors
    ```
    
     Download links
    
    1. **Text encoder (UMT5 XXL fp8)** → `models/text_encoders`
       https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors?download=true
    2. **CLIP Vision H** → `models/clip_vision`
       https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/clip_vision/clip_vision_h.safetensors?download=true
    3. **Wan 2.1 VAE** → `models/vae`
       https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/vae/wan_2.1_vae.safetensors?download=true
    4. **LightX2V I2V rank64 step-distill LoRA** → `models/loras`
       https://huggingface.co/lgylgy/Wan21_I2V_14B_lightx2v_cfg_step_distill_lora_rank64/resolve/main/Wan21_I2V_14B_lightx2v_cfg_step_distill_lora_rank64.safetensors?download=true
    5. **SAM3.1 multiplex checkpoint** → `models/checkpoints`
       https://huggingface.co/Comfy-Org/sam3.1/resolve/main/checkpoints/sam3.1_multiplex_fp16.safetensors?download=true
    6. **SCAIL-2 GGUF diffusion model** → `models/unet`
       https://huggingface.co/realrebelai/SCAIL-2_GGUF/resolve/main/SCAIL-2-Q4_K_M.gguf?download=true
    
    ---
    
     🧩 Required custom nodes
    
    - **ComfyUI-GGUF** — GGUF UNet loading
    - **ComfyUI-MultiGPU** — `UnetLoaderGGUFDisTorch2MultiGPU` (2nd-GPU weight offload)
    - **ComfyUI-KJNodes** — WanChunkFeedForward, ImageResizeKJv2, KikoPurgeVRAM, SimpleCalculatorKJ, INTConstant, GetImageRangeFromBatch, Set/Get nodes
    - **ComfyUI_Swwan** — WanSCAILToVideo, SCAIL2ColoredMask, SAM3_VideoTrack, ImageConcatMulti
    - **ComfyUI-easy-use** — forLoopStart/End, compare, ComfySwitchNode, BatchImagesNode, ColorTransfer
    - **ComfyUI-VideoHelperSuite** — VHS_LoadVideo, VHS_VideoCombine, VHS_VideoInfo
    - **ComfyUI-Resolution-Master** — ResolutionMaster
    - **rgthree-comfy** — Any Switch (Model Switcher), Display Int
    
    All of these are installable through **ComfyUI-Manager** ("Install Missing Custom Nodes").
    
    ---
    
     ⚙️ Requirements & performance
    
    - A recent ComfyUI build with **Wan 2.1 / SCAIL-2** support.
    - **~16 GB VRAM** recommended for the main GPU.
    - For MultiGPU offload: a **second GPU with ≥ 11 GB free** VRAM.
    - **Render time scales with clip length** — each window is a full diffusion pass. A ~500-frame clip runs roughly 7 windows. Use `select_every_nth` or fewer steps to trade quality/length for speed.
    
    ---
    
     🗂️ Workflow layout
    
    Nodes are organized into color-coded groups for clarity:
    
    **INPUTS** (image & video) · **MODELS** (diffusion / VAE / CLIP / sampler) · **PROMPTS** · **PREPROCESS** (resolution / pose resize / CLIP vision) · **MASK & TRACKING (SAM3)** · **CHUNK 1** (first window) · **LOOP MATH** (window / count) · **LOOP BODY** (chunk-2 generation & accumulation) · **OUTPUT** (final video) · **COMPARISON OUTPUT** (side-by-side)
    
    ---
    
     📺 Tutorial
    
    Watch how to use this workflow:
    https://www.youtube.com/@AiMotionStudio
    
    ---
    
     📝 Notes & tips
    
    - Output only appears when the full loop completes — longer clips take longer before you see anything. That's expected.
    - Keep exactly one model loader active (single- vs. dual-GPU).
    - If you hit a system-RAM error on very long/high-res inputs, keep `VHS_LoadVideo` `custom_width = 480` (already set) and/or raise `select_every_nth`.
    - Credits: built on Wan 2.1 SCAIL-2, LightX2V distill LoRA, SAM3.1, and the open-source ComfyUI custom nodes listed above.
    

    Description

    Version 2.0

    FAQ

    Comments (21)

    lidianeporto9248Jun 20, 2026
    CivitAI

    It works, but It is a time/resource-consuming workflow. Definitely a Studio-grade workflow for high-fidelity animation. Do not recommend for those who want to animate simple things (like cloning tik-tok dancing). in my case the Wan 2.2 Animate is still the best option.

    AIMotionStudio
    Author
    Jun 20, 2026

    the quality of the animaton here is way better than Wan 2.2, you don't even need to up-res the quality here.

    stavros247Jun 20, 2026
    CivitAI

    Not quick, first gen running now, multi-gpu runs like crap on my rig, I had to disable all gpus but one, a 3090, to get it to load models properly. I'll repsond with quality of the Q8 when it actually finishes.

    stavros247Jun 20, 2026

    didn't realize the source was 33sec, 54min on a 3090! 6/10 sharpness ( probably skill issue, bad source image ) , motion and audio is perfect. will use this more

    AIMotionStudio
    Author
    Jun 21, 2026

    @stavros247 the Version 3.0 is much better for people using just 1 GPU, i have updated the node to be much faster on 1GPU, however you still need to download the correct GGUF that can fit your specific GPU and also has some VRAM room left to handle other tasks

    stavros247Jun 21, 2026

    @AIMotionStudio works great, i had to source very good ref photos. thansk for the workflow! I swapped out the ggufs for stock loading with kj nodes for sage2, fp16 works great on a 3090.

    jellywing414280Jun 21, 2026
    CivitAI

    I want to keep the outfit from the original video and only change the character. Is that possible?

    AIMotionStudio
    Author
    Jul 12, 2026

    tyr to first generate the reference image with the outfit first

    stewi0001Jun 22, 2026
    CivitAI

    Does not appear to work on non-human characters. To be more specific, the results ended up with a human head. Also, these were quick trials.

    AIMotionStudio
    Author
    Jul 12, 2026

    i haven't try with non-human characters, i think you used use an image where all the 2D full body can be seen

    drfaker911219Jun 23, 2026· 1 reaction
    CivitAI

    gave it a try. unfortunately the background of the image is used instead of the video's. usually if replacement node is set to true it would be fine but in the case of this wf, it looks also set to true but it is greyed out. Unsure what to make of it

    iljalehn5993Jun 25, 2026

    just use nano to create a pic matching videos background

    drfaker911219Jun 26, 2026

    @iljalehn5993 extra annoying step. I'd rather use a different workflow that does this

    AIMotionStudio
    Author
    Jul 12, 2026

    try out version 4.0

    ProteiniqueJul 15, 2026

    @drfaker911219  hello mate, did you figure how to keep the background from the reference image instead of the video background ?

    llcjogs866Jun 28, 2026· 1 reaction
    CivitAI

    People just complain about anything and demand about everything these days. Rude's the new normal... Remember when things used to be, you know, different? x_x
    Thanks for the workflow. I wanted to upload the result here but I guess this is a different function

    AIMotionStudio
    Author
    Jul 12, 2026

    try with version 4.0

    popestmasterJun 29, 2026
    CivitAI

    Waiting for the multi-character version

    AIMotionStudio
    Author
    Jul 12, 2026

    try with version 4.0, also you most enter the number of characters in the image in the sam2 max object node!

    PdidiJul 7, 2026
    CivitAI

    im new to this motion transfer WFs, any video suggestions on how to actually tweak these nodes or what is the best way to approach these WFs with your reference image? cuz I'm getting mutations, attached body parts and all other stuff

    AIMotionStudio
    Author
    Jul 12, 2026

    try the version4.0 just released it solved most issues.

    Workflows
    Wan Video 2.2 I2V-A14B

    Details

    Downloads
    180
    Platform
    CivitAI
    Platform Status
    Available
    Created
    6/19/2026
    Updated
    8/11/2026
    Deleted
    -

    Files

    scail2GGUFMOTIONTRANSFER_v20.zip

    Mirrors