# Score to Screen — One prompt → one song → a 1–4 scene music video (YuE2 × MiniMax H3)
Write one prompt and pick your images. You get an original song made with YuE2 and a music video of 1 to 4 scenes made with MiniMax H3. The singer lip-syncs, the band plays along with the music, and every scene change lands exactly on a bar line.
v4.0 Simple: only 5 switches. Everything else is automatic, or set to the values that worked best.
---
## ⚠ CUSTOM NODES ARE NOT IN COMFYUI MANAGER
21 of the nodes come from 4 custom node packages in custom_nodes_en_v4.0.zip (included in the download). Manager can't find them; that is normal. Install them by hand:
1. Close ComfyUI.
2. Unzip, then copy the 4 folders directly into ComfyUI/custom_nodes/:
ComfyUI-PromptSplitter / ComfyUI-ABCTiming / ComfyUI-ABCScoreIO / ComfyUI-ABCInstrumentalize
✅ ComfyUI/custom_nodes/ComfyUI-PromptSplitter/__init__.py
❌ ComfyUI/custom_nodes/custom_nodes_en_v4.0/ComfyUI-PromptSplitter/... (one folder too deep, the nodes stay red)
3. Install demucs into ComfyUI's Python:
portable: .\python_embeded\python.exe -m pip install demucs
4. "Image Saver Save Video" is ComfyUI-Image-Saver; install it from Manager.
5. Restart ComfyUI.
Still red? Look for "IMPORT FAILED" in the console at startup and paste that error in the comments.
---
## 🎬 WHAT IT DOES
- One prompt, two models: a local LLM (LM Studio) splits your prompt into YuE2 style tags and lyrics, plus an H3 video prompt for each scene
- One song, 1–4 scenes: YuE2 writes a single song. H3 renders each scene from its own image, and the scenes are joined while the song plays straight through (no audible seam)
- Choose the scene count each run: 1 scene this time, 4 the next. Unused scenes don't run at all
- The score drives the picture: YuE2 also writes an ABC sheet-music plan. The workflow reads its tempo, time signature and chords, then turns them into timed cues in each scene's H3 prompt ("At 4s the vocals come in, and at the same moment…")
- Bar-accurate cuts: every scene change snaps to the start of a bar
- Band performance: write who plays what ("the white-haired woman plays the drums") and each player gets cues
- 🥁 Drum sync: an on-screen drummer follows the song's real drums. It tries STRUM (optional AI transcription) first, then the Demucs drum stem, then the score
- Drum cues are written in a way H3 can actually draw: fills as one sweep across the kit, big crashes, a few strong accents
- 🎹 Piano sync: an on-screen pianist follows the song's real key strikes (STRUM + Basic Pitch, optional)
- Singing on screen: auto / yes / no. A song can have vocals while nobody on screen lip-syncs
- Instrumentals: write "Instrumental." The workflow then:
- removes the lyrics
- moves the vocal melody to an instrument in the score
- keeps everyone's mouth closed
- Skips the intro automatically: vocal onset detection (Demucs) starts the clip just before the singing, and leaves enough song for all your scenes
- Use your own songs:
- load a song you made earlier; its .abc score is saved next to it and loaded automatically
- ▶ Start at lets you begin from the chorus, or any second you choose
- A/V sync: H3's motion lags the sound by about 0.2 s, so the final audio is shifted to match. You can adjust this
- Run log inside the workflow:
- 📺 Runs so far adds one line per run: video file name, seeds, steps, memo
- queue several runs, walk away, and check later which video came from which seed
- the log is also saved as a CSV
- Civitai-ready output: the MP4 is saved with metadata (the prompts of every scene used, YuE2 style and lyrics, H3 + YuE2 model hashes)
## ⚙️ HOW IT FLOWS
1. Prompt Splitter (LLM): makes the YuE2 style and lyrics, plus one H3 prompt per scene
2. YuE2: writes the ABC score, then sings the song (skipped when you pick a pre-made song)
3. Vocal onset detection: trims the song, then cuts the scenes on bar lines
4. ABC Timing: adds tempo, chord and band cues to each scene's H3 prompt
5. H3 ref2va: renders scene 1…N, one per image. The scenes are joined, the song is laid over them, and the video is saved
The internals are packed into subgraphs. Day to day, you only touch the input panel and the 🎛 switches.
## 🎛 THE 5 SWITCHES
- ⚡ Turbo: fast drafts (off = 20 steps)
- 🎤 On-screen singing: auto / yes / no
- 🥁 Drum sync: auto / off (piano sync follows this switch too)
- 🎵 Song delay: seconds of silence at the start of the video
- 🎬 Scene count: 1–4
Seeds (video, vocals, score) can be locked separately in the 🎲 Seed details group. Rarely used settings live in 🔧 Details:
- bar snapping
- longest scene length
- drum cue style
- A/V sync
- song volume curve
## 📦 REQUIREMENTS
Custom nodes (my own, included in the download; not in ComfyUI Manager, see ⚠ above):
- ComfyUI-PromptSplitter: prompt split and vocal onset. Run pip install demucs in ComfyUI's Python
- ComfyUI-ABCTiming: score → rhythm, band, drum and piano cues
- ComfyUI-ABCInstrumentalize: instrumental score and bar-snapped scene cuts (1–4 scenes)
- ComfyUI-ABCScoreIO: switches, saving and loading scores, pre-made songs, run log, A/V sync, scene pictures
Other custom nodes:
- ComfyUI-Image-Saver (final save with Civitai metadata; install it from Manager)
Models:
- diffusion_models: minimax_h3_ref2va_pruned_int8_convrot
- loras: minimax_h3_ref2v_turbo_4step (⚡ Turbo)
- text_encoders: qwen3vl_32b_minimax_h3_nvfp4_awq
- vae: minimax_h3_video_vae_fp16, minimax_h3_audio_vae_fp32
- checkpoints: yue2_3b_bf16 or yue2_3b_int8_convrot
LLM: LM Studio (or another OpenAI-compatible server). Load a vision model (Qwen2.5-VL, Qwen3-VL, Gemma 3…) so the LLM can see your images. A text-only model works too, but it can't look at the pictures.
Optional: STRUM, for more precise drum and piano sync. A setup guide is included. Without it, the workflow uses Demucs or the score.
## 🚀 QUICK START
1. Copy the four custom node folders directly into ComfyUI/custom_nodes (see ⚠ above), install demucs and ComfyUI-Image-Saver, then restart ComfyUI
2. Start LM Studio's server (default http://127.0.0.1:1234/v1)
3. Write scene 1's text: the scene AND the song (mood, genre, vocals or "Instrumental.")
4. Write the text for the other scenes: the scene only
5. Choose your images and the seconds for each scene. Set the 🎬 scene count; for 3–4 scenes, also fill in the 🎛 ①b group
6. Run. Check the 📺 previews: the scene cuts, the detected start, the prompts H3 got, and Runs so far
## 💡 TIPS
- Prompts can be English or Japanese. Add "Lyrics in English." to be sure of the lyric language
- "Instrumental break" or "instrumental solo" in a song with vocals is fine; it stays a song
- For a band, always write who plays what
- Turbo LoRA ON for drafts, OFF for final renders
- Scene length: the default limit is 15 s (max_scene_sec in 🔧 Details); up to 24 s has worked
- For a long song, 4 shorter scenes usually look better than 2 long ones. H3 runs once per scene, so 4 scenes take about 4× as long
- Fine finger work is still hard for H3: avoid long close-ups of hands
- Resolution: multiples of 32, short side up to 768 (1024×768 recommended)
- English and Japanese versions of the workflow are both included
## 🗂 ALSO INCLUDED: v2.34 Full / Legacy
The older "every option" workflow is still in the download and uses the same custom nodes. Use it if you need the modes the Simple line dropped:
- H3-made music (fl2va)
- song only
- sound effects only
- frozen audio
- 17 switches with fine settings
---
Version history
- v4.0: choose 1–4 scenes (🎬 scene count switch)
- v3.8: drum cues H3 can draw
- v3.7: piano sync
- v3.6: A/V sync
- v3.4 / v3.5: run log in the workflow
- v3.2: ▶ Start at for pre-made songs
- v3.0: the new Simple line (4 switches)
- v2.4: first public release
Description
Score to Screen 2-scene v2.9 (2026-10-03)
Changes since v2.6:
Image 1 no longer ignored (v2.7)
The scene images are now sent to the prompt-split LLM. With a vision model in LM Studio (e.g. Qwen3-VL Instruct), the LLM looks at each picture and describes the people, clothes and place from it.
With a text-only model, the split still works without the image. Places, clothes and hair the LLM invents are dropped, and the prompt keeps the setting "exactly as in <Picture 1>".
The camera always opens on the exact framing of the reference image.
Music cues from the score are now added inside the LLM's shots, not as extra shots. You get one clean timeline with no overlapping shots or duplicate "starts playing" lines.
Song only mode (v2.8)
New 🎬 Song only switch: makes and saves only the YuE2 song and its ABC score. H3 does not run at all.
Pick the song you like later with 🎵 Use a pre-made song, then make the video.
Models in one place (v2.9)
H3 models, Turbo LoRAs, text encoder, VAEs and the YuE2 checkpoint now sit in a 📦 Models group at the far right, out of the subgraphs.
H3 and YuE2 are chosen with 📦 pickers. The same choice sets both the loaded model and the model name in the Civitai metadata, so they always match.
Also
English edition of the workflow included (score2screen_2scene_v2.9_en.json).
Required custom nodes (included, please update)
PromptSplitter v2.0 / ABCTiming v2.4 / ABCScoreIO v1.4
FAQ
Comments (6)
This workflow is currently published for testing purposes.
hello, ok . Could you please publish an exemple ?
Thank you. I have added an example.
Thanks. This is great.
@sdktertiaire2 Thank you. While this workflow still has some issues and is essentially a test release, I hope it serves as a useful reference.
@koyan0418439 thanks
