Testing Wan2 S2V on RTX4090: Render Times, Quality, and MultiTalk Comparison

I’ve been testing the new Wan2 S2V workflow on an RTX4090 (14B_fp8). Here’s a breakdown of the results, render times, and limits I’ve found, plus a direct comparison with MultiTalk. Spoiler: S2V shows promise with audio reactivity, but still struggles with FPS and long-term consistency.
Tests #Wan2 #S2V – 13s in 4/5 (512x672 – 208f @16fps)
Model: 14B_fp8 – RTX4090 24GB VRAM + 64GB RAM – SageAttention – Same Seed
Prompt: a man is passionately singing and playing piano of an emotional song, lens flare
The workflow provided by ComfyUI is a real mess: it has to be significantly adjusted each time depending on video length and content.
It’s also difficult to automate some calculations (77-frame batches): since audio almost never matches perfectly, the last batch rarely hits 77 frames. But adjusting just that last batch adds no real value.
I still added a few Sage / Torch optimizations.
Test results:
Lora / 4 steps / Cfg 1 → 120s (5.91s/it)
LipSync OK, but little expression and very few movements.Lora Strength x1.5 / 4 steps / Cfg 1 → 87s (5.61s/it)
Lipsync OK. More movements, but way too frantic — almost Parkinson-like after a few seconds. Image quickly gets “burned” (contours).Lora Strength x0.5 / 8 steps / Cfg 2.5 → 414s (16.78s/it)
Good compromise: solid lipsync, decent movements, though a bit static at the piano.No Lora / 20 steps / Cfg 6 → 15 minutes
Good lipsync, good expressions, but not much difference from version 3 for a much longer render time. (Funny detail: it rains inside his mouth).
👉 Key notes:
Using a Lora reduces the required steps and speeds up rendering.
But it comes at the cost of movement, which becomes more static.
S2V is interesting because it reacts more to the input audio (soft or loud parts), with characters adapting their gestures slightly to the sound.
👉 The best result is still version 3, with a refined prompt:
A man passionately sings and plays a poignant song on the piano. His hands move from left to right across the piano keys. lens flare
→ Solid facial expressions, accurate lipsync, good balance between quality and render time.
Weak points:
On longer exports (55s / 874 frames), consistency collapses: lipsync disappears and the image degrades fast.
S2V: 11 minutes (874 frames @16fps) with Lora, 4 steps, Cfg 1 → impossible to render correctly with other parameters (including version 3 above).
MultiTalk: 12 minutes (1365 frames @25fps) for a stable result, precise lipsync and consistent image.
Compared to #MultiTalk / InfiniteTalk, the gap is clear:
MultiTalk runs at 25 FPS vs S2V locked at 16 FPS.
More stability and precise lipsync (S2V loses 10 frames every second).
Much more consistent, especially with singing.
Downside: MultiTalk sometimes produces overly frantic and exaggerated movements. Still, it remains clearly ahead of S2V at this stage.
Conclusion: S2V is promising, especially in how it reacts to audio, but it’s still too limited by FPS and lacks long-term consistency. MultiTalk remains the leader, even if it needs some smoothing on extreme movements.