Ready-made ComfyUI workflows for Verboa Image 1.0 (ERNIE-Image fine-tune, 8 steps, photoreal, adults only). Drag a .json onto ComfyUI and press Run.
Included
verboa-image-1.0-hires.json: fp8 safetensors (RTX 40/50) with a built-in 2.3 MP hires pass (upscale, then a second 8-step sampler at denoise 0.42).
verboa-image-1.0-gguf-hires.json: the same hires graph for the GGUF files (needs the ComfyUI-GGUF node).
verboa-image-1.0.json and verboa-image-1.0-gguf.json: the plain 1 MP originals.
Settings baked in
8 steps, CFG 2.0, euler + simple, ModelSamplingSD3 shift 4.0, empty negative prompt, FLUX.2 VAE, Ministral 3B text encoder (CLIP type flux2).
Hires pass: ImageScaleToTotalPixels 2.3 MP (steps of 16), VAE encode, 8 steps at denoise 0.42 with the same prompt. About 60 seconds per image on a laptop RTX 5090.
You also need
The model: Verboa Image (fp8 / bf16 or GGUF Q8_0) from the Verboa Image model page; text encoder ministral-3-3b and VAE flux2-vae from Comfy-Org/ERNIE-Image.
For GGUF: the ComfyUI-GGUF custom node (Unet Loader (GGUF)).
18+ only. Fictional adults only: no real people or look-alikes.
Description
Verboa Video: the video stage of our pipeline as two ComfyUI workflows. verboa-h3-from-text.json draws the first frame with Verboa Image 1.0 and animates it with MiniMax H3, with sound; verboa-h3-i2v.json animates any picture. Unzip anywhere, drag a .json into ComfyUI (tested on 0.37.4), put the five MiniMax H3 files from Comfy-Org/MiniMax-H3 in their models folders (list in the README). An 8 s clip at 512x896 takes about 110 s on a 24 GB card. Dedicated page: Verboa Video. The workflows need the MiniMax H3 weights, under MiniMax's license; mark what you post as AI-generated.