CivArchive
    One prompt → one song → a two-scene music video (YuE2 × MiniMax H3) - v3.6

    # Score to Screen — One prompt → one song → a 1–4 scene music video (YuE2 × MiniMax H3)

    Write one prompt and pick your images. You get an original song made with YuE2 and a music video of 1 to 4 scenes made with MiniMax H3. The singer lip-syncs, the band plays along with the music, and every scene change lands exactly on a bar line.

    v4.0 Simple: only 5 switches. Everything else is automatic, or set to the values that worked best.

    ---

    ## ⚠ CUSTOM NODES ARE NOT IN COMFYUI MANAGER

    21 of the nodes come from 4 custom node packages in custom_nodes_en_v4.0.zip (included in the download). Manager can't find them; that is normal. Install them by hand:

    1. Close ComfyUI.

    2. Unzip, then copy the 4 folders directly into ComfyUI/custom_nodes/:

    ComfyUI-PromptSplitter / ComfyUI-ABCTiming / ComfyUI-ABCScoreIO / ComfyUI-ABCInstrumentalize

    ✅ ComfyUI/custom_nodes/ComfyUI-PromptSplitter/__init__.py

    ❌ ComfyUI/custom_nodes/custom_nodes_en_v4.0/ComfyUI-PromptSplitter/... (one folder too deep, the nodes stay red)

    3. Install demucs into ComfyUI's Python:

    portable: .\python_embeded\python.exe -m pip install demucs

    4. "Image Saver Save Video" is ComfyUI-Image-Saver; install it from Manager.

    5. Restart ComfyUI.

    Still red? Look for "IMPORT FAILED" in the console at startup and paste that error in the comments.

    ---

    ## 🎬 WHAT IT DOES

    - One prompt, two models: a local LLM (LM Studio) splits your prompt into YuE2 style tags and lyrics, plus an H3 video prompt for each scene

    - One song, 1–4 scenes: YuE2 writes a single song. H3 renders each scene from its own image, and the scenes are joined while the song plays straight through (no audible seam)

    - Choose the scene count each run: 1 scene this time, 4 the next. Unused scenes don't run at all

    - The score drives the picture: YuE2 also writes an ABC sheet-music plan. The workflow reads its tempo, time signature and chords, then turns them into timed cues in each scene's H3 prompt ("At 4s the vocals come in, and at the same moment…")

    - Bar-accurate cuts: every scene change snaps to the start of a bar

    - Band performance: write who plays what ("the white-haired woman plays the drums") and each player gets cues

    - 🥁 Drum sync: an on-screen drummer follows the song's real drums. It tries STRUM (optional AI transcription) first, then the Demucs drum stem, then the score

    - Drum cues are written in a way H3 can actually draw: fills as one sweep across the kit, big crashes, a few strong accents

    - 🎹 Piano sync: an on-screen pianist follows the song's real key strikes (STRUM + Basic Pitch, optional)

    - Singing on screen: auto / yes / no. A song can have vocals while nobody on screen lip-syncs

    - Instrumentals: write "Instrumental." The workflow then:

    - removes the lyrics

    - moves the vocal melody to an instrument in the score

    - keeps everyone's mouth closed

    - Skips the intro automatically: vocal onset detection (Demucs) starts the clip just before the singing, and leaves enough song for all your scenes

    - Use your own songs:

    - load a song you made earlier; its .abc score is saved next to it and loaded automatically

    - ▶ Start at lets you begin from the chorus, or any second you choose

    - A/V sync: H3's motion lags the sound by about 0.2 s, so the final audio is shifted to match. You can adjust this

    - Run log inside the workflow:

    - 📺 Runs so far adds one line per run: video file name, seeds, steps, memo

    - queue several runs, walk away, and check later which video came from which seed

    - the log is also saved as a CSV

    - Civitai-ready output: the MP4 is saved with metadata (the prompts of every scene used, YuE2 style and lyrics, H3 + YuE2 model hashes)

    ## ⚙️ HOW IT FLOWS

    1. Prompt Splitter (LLM): makes the YuE2 style and lyrics, plus one H3 prompt per scene

    2. YuE2: writes the ABC score, then sings the song (skipped when you pick a pre-made song)

    3. Vocal onset detection: trims the song, then cuts the scenes on bar lines

    4. ABC Timing: adds tempo, chord and band cues to each scene's H3 prompt

    5. H3 ref2va: renders scene 1…N, one per image. The scenes are joined, the song is laid over them, and the video is saved

    The internals are packed into subgraphs. Day to day, you only touch the input panel and the 🎛 switches.

    ## 🎛 THE 5 SWITCHES

    - ⚡ Turbo: fast drafts (off = 20 steps)

    - 🎤 On-screen singing: auto / yes / no

    - 🥁 Drum sync: auto / off (piano sync follows this switch too)

    - 🎵 Song delay: seconds of silence at the start of the video

    - 🎬 Scene count: 1–4

    Seeds (video, vocals, score) can be locked separately in the 🎲 Seed details group. Rarely used settings live in 🔧 Details:

    - bar snapping

    - longest scene length

    - drum cue style

    - A/V sync

    - song volume curve

    ## 📦 REQUIREMENTS

    Custom nodes (my own, included in the download; not in ComfyUI Manager, see ⚠ above):

    - ComfyUI-PromptSplitter: prompt split and vocal onset. Run pip install demucs in ComfyUI's Python

    - ComfyUI-ABCTiming: score → rhythm, band, drum and piano cues

    - ComfyUI-ABCInstrumentalize: instrumental score and bar-snapped scene cuts (1–4 scenes)

    - ComfyUI-ABCScoreIO: switches, saving and loading scores, pre-made songs, run log, A/V sync, scene pictures

    Other custom nodes:

    - ComfyUI-Image-Saver (final save with Civitai metadata; install it from Manager)

    Models:

    - diffusion_models: minimax_h3_ref2va_pruned_int8_convrot

    - loras: minimax_h3_ref2v_turbo_4step (⚡ Turbo)

    - text_encoders: qwen3vl_32b_minimax_h3_nvfp4_awq

    - vae: minimax_h3_video_vae_fp16, minimax_h3_audio_vae_fp32

    - checkpoints: yue2_3b_bf16 or yue2_3b_int8_convrot

    LLM: LM Studio (or another OpenAI-compatible server). Load a vision model (Qwen2.5-VL, Qwen3-VL, Gemma 3…) so the LLM can see your images. A text-only model works too, but it can't look at the pictures.

    Optional: STRUM, for more precise drum and piano sync. A setup guide is included. Without it, the workflow uses Demucs or the score.

    ## 🚀 QUICK START

    1. Copy the four custom node folders directly into ComfyUI/custom_nodes (see ⚠ above), install demucs and ComfyUI-Image-Saver, then restart ComfyUI

    2. Start LM Studio's server (default http://127.0.0.1:1234/v1)

    3. Write scene 1's text: the scene AND the song (mood, genre, vocals or "Instrumental.")

    4. Write the text for the other scenes: the scene only

    5. Choose your images and the seconds for each scene. Set the 🎬 scene count; for 3–4 scenes, also fill in the 🎛 ①b group

    6. Run. Check the 📺 previews: the scene cuts, the detected start, the prompts H3 got, and Runs so far

    ## 💡 TIPS

    - Prompts can be English or Japanese. Add "Lyrics in English." to be sure of the lyric language

    - "Instrumental break" or "instrumental solo" in a song with vocals is fine; it stays a song

    - For a band, always write who plays what

    - Turbo LoRA ON for drafts, OFF for final renders

    - Scene length: the default limit is 15 s (max_scene_sec in 🔧 Details); up to 24 s has worked

    - For a long song, 4 shorter scenes usually look better than 2 long ones. H3 runs once per scene, so 4 scenes take about 4× as long

    - Fine finger work is still hard for H3: avoid long close-ups of hands

    - Resolution: multiples of 32, short side up to 768 (1024×768 recommended)

    - English and Japanese versions of the workflow are both included

    ## 🗂 ALSO INCLUDED: v2.34 Full / Legacy

    The older "every option" workflow is still in the download and uses the same custom nodes. Use it if you need the modes the Simple line dropped:

    - H3-made music (fl2va)

    - song only

    - sound effects only

    - frozen audio

    - 17 switches with fine settings

    ---

    Version history

    - v4.0: choose 1–4 scenes (🎬 scene count switch)

    - v3.8: drum cues H3 can draw

    - v3.7: piano sync

    - v3.6: A/V sync

    - v3.4 / v3.5: run log in the workflow

    - v3.2: ▶ Start at for pre-made songs

    - v3.0: the new Simple line (4 switches)

    - v2.4: first public release

    Description

    # Score to Screen v3.6 Simple — Changelog (from v2.19)

    **One prompt → an original song + a 2-scene music video (YuE2 × MiniMax H3)**

    > ### ⚠ v3 is a new, simpler line — not just the next v2 update

    > v2.x grew into a "full control" workflow: 17 switches, several modes (H3-made music, song only, SFX only, frozen audio) and many fine settings.

    > **v3.x Simple goes the other way.** It keeps only the everyday flow — prompt → song → 2-scene video — with **just 4 switches**. Everything else is decided automatically or fixed to the values that worked best in v2.

    **v3.x Simple (this release)**

    - Aim: easy, good results with little setup

    - Switches: 4 (Turbo, on-screen singing, drum sync, song delay)

    - Modes: song + video (YuE2 song, or your pre-made song)

    - Best for: most users, first-timers, batch runs

    **v2.34 Full / Legacy**

    - Aim: every option exposed

    - Switches: 17 + detail settings

    - Modes: also H3-made music, song only, SFX only, frozen audio

    - Best for: users who want to fine-tune or need the extra modes

    **Both are included and use the same custom_nodes**, so you can switch between them at any time. If you relied on a v2-only mode, keep using **v2.34 Full / Legacy**.

    ---

    ## 🆕 New in v3.3 – v3.6

    - 🎚 **A/V sync (v3.6).** H3's motion (hands, sticks) comes about **0.2 s after** the sound it hears. H3's own audio matches the song exactly; only the motion lags.

    - The final video's sound now plays 0.2 s later to match.

    - It is taken from 0.2 s earlier in the song, so nothing goes silent. Only a song used from its very start gets 0.2 s of silence first.

    - Adjust it in 🔧 Details → 🎚 A/V sync (0 = off, negative = sound earlier). H3's reference audio, the score and the cues are unchanged.

    - 📺 **Runs so far, inside the workflow (v3.4 / v3.5).** Each run adds one line, newest first:

    - the video file name, 🎬 video seed, steps, 🎤 vocal / 🎼 score seeds, and a 📝 memo

    - queue several runs and leave them; later you can see which video came from which seeds

    - the runs are also kept in output/video/score2screen_run_log.csv (opens in Excel), so the list continues after a restart

    - 🥁 **Clearer drum-sync messages (v3.3).** 📺 Drum hits now says why there is no hit table:

    - synced with Demucs (success), no drummer on screen, failed, or off

    - 🥁 **STRUM errors explained (v3.3).**

    - When STRUM makes no MIDI, the reason and a hint come first (e.g. demucs missing in STRUM's own Python).

    - The full log is saved to output/score2screen_strum_cache/strum_last_log.txt.

    - The setup guide now installs STRUM with pip install -e ".[separation]" and has a Troubleshooting section.

    ## 🎵 v3.2: ▶ Start at for pre-made songs

    - Use your song from the chorus, or from anywhere you like.

    - auto (default): starts at the detected vocal onset, as before.

    - Seconds 62.5 or minutes:seconds 1:02.5: starts exactly there.

    - The start is moved earlier only if both scenes would run past the end of the song (shown in 📺).

    - With a score, the bar-snapped scene cut, the playing cues and the lyric timing all follow the new start.

    ## ✨ Highlights of the Simple redesign (v3.0 / v3.1)

    - 🎛 **Only 4 switches** (new Simple Board node):

    - ⚡ **Turbo**: fast generation (when off, 20 steps)

    - 🎤 **On-screen singing**: auto / yes / no

    - 🥁 **Drum sync**: auto / off

    - 🎵 **Song start delay** (seconds of silence at the start of the video)

    - 🎤 **"Song has vocals" and "the person on screen sings" are separate decisions**

    - A song can have vocals while nobody on screen lip-syncs (mouths stay closed).

    - auto: the LLM decides from your text and images. Scene 2 follows scene 1's decision.

    - The vocal / instrumental decision is made in this order:

    1. strong instrumental words

    2. vocals stated explicitly

    3. "doesn't sing" alone means instrumental

    4. otherwise the LLM decides

    - 🥁 **Drum sync auto**: the on-screen drummer follows the real drums of the song

    - Tries **STRUM** (neural drum transcription, optional), then **Demucs** drum-stem hits, then the **ABC score**.

    - Only runs when someone on screen plays the drums.

    - 🤖 **Now automatic / fixed**:

    - vocal onset detection (instrumentals start at 0 s)

    - cues for the other band members

    - first camera change no earlier than 3 s

    - extra style words for instrumentals

    - final audio is always the YuE2 song

    - 🎲 **Seeds** (video / vocals / score) are tucked into a collapsed "Seed details" group.

    - 🔧 **Rarely used settings** are grouped under "Details":

    - snapping the scene cut to bars

    - the song volume curve

    - A/V sync (v3.6)

    ## 🗂 Removed from Simple (still in v2.34 Legacy)

    - H3-made music mode (fl2va) and its model / LoRA

    - Song-only mode, SFX-only mode, audio freeze

    - Experimental 8-class drum hit classification

    - YuE2 / H3 final-audio switch

    ## 🎶 Notable improvements since v2.19

    **Pre-made songs**

    - Just pick a song and it's used automatically (YuE2 is skipped).

    - Choose where in the song to start with ▶ Start at (v3.2).

    - An .abc score with the same name is picked up automatically, even when the numbering differs.

    - 📺 shows which score was used.

    - Silence at the start of the song is measured, so no one "plays" during a silent intro.

    - Write "with vocals" or "instrumental" in scene 1's text to set the vocal type.

    **Band performance cues**

    - The splitter writes **who plays what** (e.g. "the blue-haired woman plays the guitar, the white-haired woman plays the drums").

    - Chord progressions are back in the rhythm cues.

    - Drums: hi-hat / snare / kick patterns, fills and crashes at section changes, descending tom fills.

    - Bass / keys move on chord changes; a second guitarist or drummer also gets cues.

    - Synced drum hits: kick and snare only, at most one per beat, up to 16 per shot.

    **Timing**

    - Vocal onset leaves room for both scenes.

    - The scene cut snaps to bar lines, and the onset check uses the same cut position.

    - First camera change is not before 3 s.

    - The final sound is lined up with H3's motion (A/V sync, v3.6).

    **Prompt quality**

    - Works around Civitai cutting off the first characters of the mp4 prompt.

    - No disclaimers in the H3 prompt.

    - Everyday words (car keys, oil drum…) are no longer mistaken for instruments.

    ## 📦 Requirements / How to update

    - **custom_nodes v3.6** (included zip):

    - ComfyUI-PromptSplitter **v2.10**

    - ComfyUI-ABCTiming **v2.19**

    - ComfyUI-ABCScoreIO **v1.21**

    - ComfyUI-ABCInstrumentalize **v1.3**

    - Replace your old folders with these, then restart ComfyUI. Full install steps: README.md in the zip.

    - **STRUM is optional.** Without it, drum sync falls back to Demucs or the score. Setup guide: custom_nodes/ComfyUI-ABCTiming/STRUM_setup_guide.md

    - Each scene is up to 15 s. H3 runs twice, so it takes about 2× as long as the 1-scene version.

    - Both JA and EN versions of the workflow are included.

    FAQ

    Comments (1)

    iGor777999Oct 10, 2026
    CivitAI

    выдает ошибку пишет неизвестный пакет ,хотя комфи обновил-менеджер ничего не показывает там пусто ,ни в установленных ,ни в те что установить .
    Неизвестный пакет22

    ABCInstrumentalize

    ABCSongDelay

    ABCTimingPrompt

    ABCTimingPrompt

    ABCTimingPrompt

    ABCTimingPrompt

    CheckpointLoaderByName

    CivitaiPromptFormat

    Image Saver Save Video

    PickCheckpoint

    PickDiffusionModel

    PromptSplitterLLM

    PromptSplitterLLM

    PromptSplitterLLM

    PromptSplitterLLM

    S2SDrumEventTable

    S2SDrumEventTable

    S2SSeedBoard

    S2SSimpleBoard

    SongSfxMix

    UNETLoaderByName

    VocalOnsetDetector

    Workflows
    MiniMax H3

    Details

    Downloads
    48
    Platform
    CivitAI
    Platform Status
    Available
    Created
    10/9/2026
    Updated
    10/10/2026
    Deleted
    -

    Files

    onePromptOneSongATwoScene_v36_3281800.zip

    onePromptOneSongATwoScene_v36_3281799.zip