Our own video pipeline, free: a Verboa Image still goes in, a MiniMax H3 clip with sound comes out. Two ComfyUI workflows, core nodes, 4 steps with the turbo LoRA: about 110 s for an 8 s clip at 512x896 on a 24 GB card. 18+ only.
In the zip: verboa-h3-from-text.json (Verboa Image 1.0 draws the first frame, H3 animates it, both are saved) and verboa-h3-i2v.json (animate any picture), with a README.
Set up: put the five MiniMax H3 files from Comfy-Org/MiniMax-H3 on Hugging Face in their models/ folders (list, sizes and links in the README and in the workflow's note) and drag a workflow into ComfyUI. For the first frame use Verboa Image 1.0 or any picture.
Prompting: H3 reads its own structured format (shots, time-coded beats, a soundscape line); the note in the workflow shows it with a working example. Motion and NSFW LoRAs made for H3 chain after the turbo LoRA.
What we learned making clips: the first frame decides the clip (the subject facing the camera, fully in view); describe that frame exactly in the 00:00 line; time-coded beats and handheld phone footage beat tag lists.
The workflows need the MiniMax H3 weights, which come under the MiniMax H3 Community License Agreement: it excludes use in the US, the EU, the UK and South Korea and has an acceptable use policy. Read it before you download, and mark what you post as AI-generated.
Guide with examples: Verboa Video article. Prompt guide for the image model: Verboa Image prompt guide.
Description
Two MiniMax H3 workflows with sound, plus README, NOTICE and LICENSE. Unzip anywhere and drag a .json into ComfyUI (tested on 0.37.4): verboa-h3-from-text.json (Verboa Image 1.0 draws the still, H3 animates it) or verboa-h3-i2v.json (any picture). Put the five MiniMax H3 files from Comfy-Org/MiniMax-H3 in their models folders (list in the README). Defaults: 512x896, 12 s at 24 fps, 4 steps with the turbo LoRA, and an explicit example prompt. The H3 clips on our profile are made this way. The workflows need the MiniMax H3 weights, under MiniMax's license.