This is the SD 1.5 version of Pony V6 fine-tuned at native 1024px on ~5000 hand-selected images (some of them borrowed from my Zootvision model's datasets). Each image was captioned both with Florence-2 Large "More Detailed Mode" rich captions and also Booru tags from WD-VIT-V3. Use this the same way you'd use Pony V6 SD 1.5 normally, and just uh, enjoy the pretty objectively better overall aesthetics.
Important notes:
in A111, use Clip Skip 1 (NOT 2) and in Comfy just do not use the "Clip Set Last Layer" at all, with this
DO NOT use any VAE besides the one that is baked into the checkpoint, it will fry the image
You will probably get worse results from weird "in between" resolutions like 720x1280 than you will from standard ones like 768x1024 and 832x1216
Basic positive prompt: score_9, source_whatever, rating_whatever, your tags or natural language description here.
Basic negative prompt: score_3_up, score_4_up, score_5_up, sketch, (simple background:1.2).
Recommended steps / sampler CFG: typically Euler Ancestral at CFG 7.0 with around 25 - 35 steps is a good starting place. The DPM++ SDE GPU family of samplers can also be good with this at lower CFG (4.0 - 5.0) if you're going for realism in particular.
Generating at 512x512 is NOT recommended, as this model was originally trained by AstraliteHeart at 768px, and my additional training was entirely done at 1024px.
Description
SIGNIFICANTLY strengthens source_anime with custom training. Use a combination of source_anime, 2d, anime screencap and cgi, 3d, raw, photo, realistic, photorealistic in your negative and positive to get realistic or non-realistic results as desired.
FAQ
Comments (26)
Trained it a bunch on my anime dataset. This version will give quite different results if you use source_anime, 2d and anime screencap in the positive prompt than previous versions. You can use 3d, cgi, raw, photo, realistic, photorealistic in any combination also in the positive or negative as needed.
try version for: 2.0 , 3.0 and sf 2.1
Very interesting, uploaded images with identical settings to v3 and v4 for comparison. Still need to test tags that were strengthened.
Nice! Note that "simple background" in the negative has an extreme impact on background quality, you probably want to always have it in there. I don't think "flat lineart" means anything in particular in Pony BTW, or at least I'm not aware of it. If you want 2d, you're going to want to control things with the particular tags I mentioned.
@diffusionfanatic1173 "flat lineart" was improving quality for the base Pony 1.5 (as if stabilizing sketch styles), but here it's mostly meaningless. I will post more examples later!
@dobomex761604 FYI also the prompt base I recommend in the description 100% definitely is the actual best way to start, also, like using score_9 down to score_6_up in the positive and no scores in the negative will give you worse results always than just score_9 in the positive and score_3_up, score_4_up, score_5_up in the negative. It doesn't work the same way as Pony XL.
@diffusionfanatic1173 I'm not so sure about "score_3_up, score_4_up, score_5_up in the negative", since they should be score_3, score_4, etc. - specifically for negative. Plus, that whole negative I used in the example isn't mine - I got it from somewhere here on Civitai, and it worked well on Pony 1.5 unlike many others. I essentially had to use 3 negatives to make Pony 1.5 output good results. Your models (both v3 and v4) are much better at working with negatives - more examples from PonyXL should work.
@dobomex761604 I'm not sure that the non-up versions exists for numbers that low, he never really confirmed that. Skipping six (like it not being in the positive or negative) seems to work as a "separator" though.
Added another example of how to prompt for anime in particular in V4. Don't use underscores anywhere ever outside the Pony source and rating tags. Also be sure to escape round brackets in Booru tags \(like this\)
Hey i'm making a merge with merge blocks and your model seems awesome for it since i made a 1.5x pony finetune too ; would you mind if i used merge blocks with your model?
EDIT : Just saw the permissions under the model ; merging is ok :D
Yeah go for it. I'm probably going to release at least one more version of this though, with a fair amount more training, so if you merge with 4.0 you might have to update later. What does your finetune focus on dataset-wise, btw? How many images approx? (Just asking out of curiousity).
@diffusionfanatic1173 About 2.5k pics from various known artists that are really detailed ; also I'm merging it using merge blocks with some of my previous fine tuned projects trying to preserve pony clip and integrating Nai tags so that it can be more lora compatible ; I want to achieve a really smooth anime feel without too many "heavy" details.
Probably going to do a V5 that cleans up the overall level of detail and aesthetic a bit more, with around another 600 or so handselected and well-captioned images. After that I'll probably leave this alone for a while for real lol.
I've been working on artist lora (Shikarii he makes realistic art female and futa) for pony 6 and made dataset (pics and captions). Are you interested or it's already too late?
Right now it's around 40 pics, they are not cropped to square. Watermark removed.
Pics are hires (medium quality pics were upscaled and downscaled to 2048).
If you interested I can make more captions, and fix some mistakes (some pics I saved in wrong resolution and forgot to remove watermark).
If you were already working on it you might as well keep going, I could always add it another time too
@diffusionfanatic1173 No problem, I understand.
I think my dream for a "BestPony" model would be one that works with regular prompts, without having to add so many formulas, because I tend to compare models and prompts, Pony always comes up behind if source_anime or score_9 isn't in there. Like a BetterPonyAnime that does it as if the source_anime was baked in the model and always present, and BetterPonyQuality that acts as if score_9 was always present, or something. ZootVision never needed all that, I guess I'd like a ZootVision with Pony's style, maybe it just needs a "by PonyXL" token that does it.
@Zenyth what style do you mean exactly, the inherent artstyle of Pony is kinda just like "amateur Deviantart pencil drawing" lol
Something I should note more visibly: if you were going to train a Lora on this for some reason, you NEED to do it at no lower than 768x768 overall resolution in the trainer, which is the resolution that AstraliteHeart originally finetuned base SD 1.5 on using the same dataset as PonyXL. Training at 512x512 will give you terrible results as none of the data he or I have added is bucketed that low.
So far I'm torn on clip skip: V4 works best with 1, sure, but V3 works well with both (different results, but both are good).
@diffusionfanatic1173 I've uploaded an example of clip skip 1 breaking the image on V3, just in case in helps.
@dobomex761604 The original Pony SD 1.5 is a Clip Skip 1 direct finetune of base SD 1.5. "Clip Skip 2" has nothing to do with anything other than how NovelAI trained their totally unrelated SD 1.5 anime model. All my additional training for every version of this tune of Pony SD 1.5 has been done at Clip Skip 1. TLDR this model has no NAI "DNA" whatsoever.
I don't see which one is supposed to be broken on V3. I can say that the horizon is way off kilter in your Clip Skip 2 beach one with V4, however.
@diffusionfanatic1173 The hands are merged into one on the clip skip 1 version. In any case it's easier to just leave clip skip to 1 and fix any oddities with proper parameters.
Details
Available On (1 platform)
Same model published on other platforms. May have additional downloads or version variants.





