This is the SD 1.5 version of Pony V6 fine-tuned at native 1024px on ~5000 hand-selected images (some of them borrowed from my Zootvision model's datasets). Each image was captioned both with Florence-2 Large "More Detailed Mode" rich captions and also Booru tags from WD-VIT-V3. Use this the same way you'd use Pony V6 SD 1.5 normally, and just uh, enjoy the pretty objectively better overall aesthetics.
Important notes:
in A111, use Clip Skip 1 (NOT 2) and in Comfy just do not use the "Clip Set Last Layer" at all, with this
DO NOT use any VAE besides the one that is baked into the checkpoint, it will fry the image
You will probably get worse results from weird "in between" resolutions like 720x1280 than you will from standard ones like 768x1024 and 832x1216
Basic positive prompt: score_9, source_whatever, rating_whatever, your tags or natural language description here.
Basic negative prompt: score_3_up, score_4_up, score_5_up, sketch, (simple background:1.2).
Recommended steps / sampler CFG: typically Euler Ancestral at CFG 7.0 with around 25 - 35 steps is a good starting place. The DPM++ SDE GPU family of samplers can also be good with this at lower CFG (4.0 - 5.0) if you're going for realism in particular.
Generating at 512x512 is NOT recommended, as this model was originally trained by AstraliteHeart at 768px, and my additional training was entirely done at 1024px.
Description
v5.5 + another 1000-image dataset (in preparation for the Flux one I'm gonna add for V6.0). So yeah this is still not QUITE V6.0. Definitely better than v5.5 but v6.0 should be like more of a serious step up. It will be out soon, I had to do things this way to avoid frying everything as usual.
FAQ
Comments (6)
V6 pre-release notes:
- very long, highly descriptive LLM-style positive prompts will usually produce a really nice-looking image that is at least close to what you wanted now (don't expect ZootVision-like prompt adherence though, it's just never gonna happen), and this sort of prompting works better with no score tags in the positive and literally no negative whatsoever
- the score tags are however still highly beneficial for more traditionally Ponyish terse tag-base prompting
- this is the result of me training it continuously on images that didn't use the score system at all but rather a hybrid "natural language description followed immediately by tags" captioning style over the past little while, which I've felt to be the most effective way to improve it both visually and in terms of prompt adherence without degrading the inherent base model knowledge too much
Can you share what is approximate ratio of human art to machine art in Better Pony datasets? I know that original Pony 6 (both sdxl and SD1.5 version) has strong reaction if you put "stable diffusion" into the prompt. If you use it at all.
@Cum_Miser V1 was pumped full of like ~85% actual IRL photographs before I even released it, I didn't start counterbalancing with more 2D content of any kind until a bit after. The datasets of stuff that wasn't actual photos have been relatively balanced, though.
@diffusionfanatic1173 Thanks for elaborating on the dataset, it's helpful to know.
Details
Available On (1 platform)
Same model published on other platforms. May have additional downloads or version variants.







