CivArchive
    ← All articles
    Published August 31, 2025by dodemodexter

    Testing Wan2 S2V on RTX4090: Render Times, Quality, and MultiTalk Comparison

    196 views2 reactions0 comments on CivitAI1 collected
    multitalkwans2vs2vwan22comparative studyinfinitetalk

    I’ve been testing the new Wan2 S2V workflow on an RTX4090 (14B_fp8). Here’s a breakdown of the results, render times, and limits I’ve found, plus a direct comparison with MultiTalk. Spoiler: S2V shows promise with audio reactivity, but still struggles with FPS and long-term consistency.

    Tests #Wan2 #S2V – 13s in 4/5 (512x672 – 208f @16fps)
    Model: 14B_fp8 – RTX4090 24GB VRAM + 64GB RAM – SageAttention – Same Seed

    Prompt: a man is passionately singing and playing piano of an emotional song, lens flare

    The workflow provided by ComfyUI is a real mess: it has to be significantly adjusted each time depending on video length and content.


    It’s also difficult to automate some calculations (77-frame batches): since audio almost never matches perfectly, the last batch rarely hits 77 frames. But adjusting just that last batch adds no real value.
    I still added a few Sage / Torch optimizations.

    Test results:

    1. Lora / 4 steps / Cfg 1 → 120s (5.91s/it)
      LipSync OK, but little expression and very few movements.

    2. Lora Strength x1.5 / 4 steps / Cfg 1 → 87s (5.61s/it)
      Lipsync OK. More movements, but way too frantic — almost Parkinson-like after a few seconds. Image quickly gets “burned” (contours).

    3. Lora Strength x0.5 / 8 steps / Cfg 2.5 → 414s (16.78s/it)
      Good compromise: solid lipsync, decent movements, though a bit static at the piano.

    4. No Lora / 20 steps / Cfg 6 → 15 minutes
      Good lipsync, good expressions, but not much difference from version 3 for a much longer render time. (Funny detail: it rains inside his mouth).

    👉 Key notes:

    • Using a Lora reduces the required steps and speeds up rendering.

    • But it comes at the cost of movement, which becomes more static.

    • S2V is interesting because it reacts more to the input audio (soft or loud parts), with characters adapting their gestures slightly to the sound.

    👉 The best result is still version 3, with a refined prompt:
    A man passionately sings and plays a poignant song on the piano. His hands move from left to right across the piano keys. lens flare
    → Solid facial expressions, accurate lipsync, good balance between quality and render time.

    Weak points:

    On longer exports (55s / 874 frames), consistency collapses: lipsync disappears and the image degrades fast.

    • S2V: 11 minutes (874 frames @16fps) with Lora, 4 steps, Cfg 1 → impossible to render correctly with other parameters (including version 3 above).

    • MultiTalk: 12 minutes (1365 frames @25fps) for a stable result, precise lipsync and consistent image.

    Compared to #MultiTalk / InfiniteTalk, the gap is clear:

    • MultiTalk runs at 25 FPS vs S2V locked at 16 FPS.

    • More stability and precise lipsync (S2V loses 10 frames every second).

    • Much more consistent, especially with singing.

    • Downside: MultiTalk sometimes produces overly frantic and exaggerated movements. Still, it remains clearly ahead of S2V at this stage.

    Conclusion: S2V is promising, especially in how it reacts to audio, but it’s still too limited by FPS and lacks long-term consistency. MultiTalk remains the leader, even if it needs some smoothing on extreme movements.