Image-to-video workflow for WAN 2.2 I2V (GGUF), built to run on a 12GB VRAM GPU.
Workflow
1. Generate an image with Stable Diffusion (Illustrious checkpoint)
2. Turn it into a video with this ComfyUI workflow
Required models
- Video model: DaSiWa-WAN 2.2 I2V 14B TastySin v8 | Lightspeed | GGUF (High and Low)
https://civarchive.com/models/2190659/dasiwa-wan-22-i2v-14b-tastysin-v8-or-lightspeed-or-gguf
I use Q6 High and Q6 Low on 12GB VRAM. Choose the GGUF size that fits your GPU.
- LoRA: SmoothMix Animations WAN 2.2 (High and Low)
https://civarchive.com/models/2040641/smoothmix-animations-wan-22
- Custom node for loading GGUF models (e.g. ComfyUI-GGUF)
Settings
- Resolution: 768 x 768
- Length: 81 frames / Frame rate: 32
- Steps: 5 total (3 on High, 2 on Low)
- Sampler: lcm / Scheduler: beta
- CFG: High 2.4 / Low 1.0 (adjust depending on how well the prompt works; above 3 tends to break the video)
- LoRA strength: High 0.7 / Low 1.0
Tips
- If the video barely moves from the original image, lower the LoRA strength on the High model.
- If a specific motion looks off, add a LoRA made for that motion.
- One video takes about 15 minutes on my 12GB GPU.