This ComfyUI tutorial workflow demonstrates Wan3.0’s reference-to-video generation: starting from two images, it produces a lifelike, up-to-30-second video that preserves character identity and visual style. Two LoadImage nodes bring in your references (for example, a clean subject image and a context/pose or style image). These feed directly into the Wan3ReferenceToVideoApi node, which sends the request to the Wan3.0 service and returns a ready-to-save video. A SaveVideo node then writes the final clip to disk as an MP4.
Technically, the pipeline is lean and reliable: the heavy lifting happens remotely via the Wan3ReferenceToVideoApi node, so your local GPU is not a bottleneck. The MarkdownNote node in the graph includes inline tips and reminders about inputs and expected outputs. Because the generator is reference-driven, the quality and composition of your two images strongly influence identity lock, motion coherence, and background fidelity. Use this when you want consistent character animation from stills, quick scene mocks from storyboards, or product visuals derived from brand imagery.
Description
Wan3.0: Reference to Video
Generate a video from two reference images using Wan3.0's reference-to-video capability, which produces up to 30 seconds of footage with character consistency and lifelike visuals. Ideal for creating consistent character animations, scene replication from visual references, and producing audiovisual content from static imagery.