Verboa Image: uncensored photoreal NSFW model for adults. ERNIE-Image fine-tune that finishes in 8 steps, runs on a 16 GB GPU or a Mac, and understands plain-language prompts. Free.
Settings that work
8 steps, CFG 2.0, euler + simple scheduler, ModelSamplingSD3 shift 4.0, empty negative prompt.
About 1 MP: 848x1264, 1024x1024, 1264x848. For a sharper 2.3 MP image: upscale 1.47x, then re-run 8 steps at denoise 0.4 with the same prompt and seed.
Files: fp8 (RTX 40/50, fastest) and bf16 (balanced). GGUF Q8/Q4, NVFP4 and int8 are on Hugging Face and arriving here as new versions.
Copy-paste prompts (with seeds)
a photo of a 27 year old Latina woman with platinum blonde hair and a busty figure and dimples, naked, kneeling with her hands on her thighs, breasts forward, between rows of vines in a vineyard, soft evening light. (seed 786785433, 848x1264)
an editorial photo of a 26 year old Latina woman with silver-blonde hair and a curvy figure and a belly-button piercing, naked, on all fours looking back over her shoulder at the camera, in a cedar sauna, amber light and beads of sweat on her skin. (seed 952905615, 848x1264)
Write it the way you would say it: the person (hair, build, age), what she wears or does not, what she is doing, then setting and light. Say what stays on, and use inline negations such as "no tattoos".
An 8B photorealistic text-to-image model for adults, fine-tuned from Baidu's ERNIE-Image. It finishes an image in 8 steps at guidance 2.
18+ only. Released under the CreativeML Open RAIL++-M License with one added use restriction: no sexually explicit, nude or intimate images of a real, identifiable person without that person's explicit prior consent. The use restrictions travel with every copy, merge and derivative. Read the license.
What it is
A full fine-tune of ERNIE-Image on about 300,000 captioned images: 250,000 professional adult studio photographs and 50,000 openly available general images, trained for 11,338 updates. ERNIE-Image-Turbo's few-step ability was transplanted onto it, so 8 steps are enough. Every training caption was written as the prompt a particular person would type, so plain language works, whether slang, casual, formal or clinical.
The sample images were rendered at 8 steps, guidance 2 and 848×1264, each picked from four seeds and shown without retouching. Every face in them was screened at 21 or older.
Setup in ComfyUI
This file is the diffusion model only. ERNIE-Image is supported natively in ComfyUI, so the image model needs no custom nodes.
models/diffusion_models/: this file (fp8 for NVIDIA, bf16 for Mac)models/text_encoders/:ministral-3-3b.safetensorsfrom Comfy-Org/ERNIE-Imagemodels/vae/:flux2-vae.safetensorsfrom Comfy-Org/ERNIE-Image
Ready-made workflows for NVIDIA, Mac and GGUF are on GitHub. If you build your own graph:
8 steps and CFG 2.0, with the euler sampler and the
simplescheduler. Don't drop CFG to 1: faces suffer.ModelSamplingSD3 at shift 4.0. ComfyUI's default shift for ERNIE is 3.0.
An empty CLIPTextEncode as the negative, not ConditioningZeroOut.
About 1 MP: 848×1264, 1024×1024, 1264×848, 768×1376, 1376×768, 896×1200 or 1200×896.
Every other format is on Hugging Face: fp32, int8 and NVFP4 single files, GGUF from F16 down to Q4_0, MLX for mflux on a Mac, and NF4 for diffusers.
Prompting
Write what you want the way you would say it. A sentence of 15 to 40 words is the sweet spot. The core of a prompt is the person (hair, build, age), what they are wearing or not, and what they are doing; setting, light and camera are extras. State an adult age ("in her twenties", "a man in his thirties"). Negations work ("no tattoos"). Tag lists like "masterpiece, best quality" were not trained, and there is no prompt enhancer.
Prompt Writer
Rather not write the prompt? Verboa Prompt Writer is a small companion model trained on the same captions that writes prompts for this model in ComfyUI, with one custom node: Let Me Choose (set people, age, place, how explicit, voice and length, or leave them open), Surprise Me (it picks for you) or Write My Own (your own prompt, through the same workflow). Each run writes a new prompt and renders it. The node, the files and two ready workflows are on its page.
Safety
The training data is commercial studio photography only, age-screened, and all 299,996 training images passed a CSAM hash scan with zero matches.
A safety edit trained into the weights turns a request for a child into an adult. On 200 held-out child prompts, images with a face read as under 18 fell from 80.5% to 6.0%, and 14 or under from 61.0% to 2.0%. It makes a child far less likely, not impossible, and the license forbids trying.
The full record is in the safety report. Report misuse or a safety problem to [email protected].
Limitations
It is a photographic model. For drawn styles, ERNIE-Image itself is better.
Male anatomy is slightly less reliable than female.
Partnered scenes can drift: a neighboring act, or an extra person. Several seeds help.
It may occasionally draw a faint signature in a corner.
It was trained at about 1 MP, and it is a few-step model, so LoRAs trained on it can behave oddly.
Links
Hugging Face: every format and the full model card
GitHub: workflows, an inference script, the Prompt Writer's node, and issues
Verboa Prompt Writer: writes prompts for this model in ComfyUI
Anything else: [email protected]
Description
GGUF Q8_0 of Verboa Image 1.0 for the ComfyUI-GGUF custom node (Unet Loader (GGUF)). Near-lossless, about 8.1 GB, fits 12-16 GB cards. Same settings as every other file: 8 steps, CFG 2.0, euler + simple, ModelSamplingSD3 shift 4.0, empty negative prompt.


