I'm working on a small AI Discord Community you can join here: https://discord.gg/ehhMrD4PQT
My Very First VAE Merge, which is found in one of my models on my old account. Feel free to test and/or use them. V3 is what I originally wanted the VAE to be, a VAE that adds subtle contrast/saturation/brightness.
This will not fix an already overfitted/overbaked image as it's a VAE. Also the VAEs may effect images a bit too much for you personally, especially if you're doing hires/upscale/etc with IMG2IMG, but if you like the those images go for it. Works best for TXT2IMG. I personally like the VAEs for TXT2IMG mainly, and IM2IMG if the image is washed.
Also VAE has a tendency to fix some minor background stuff sometimes, due to it properly doing that artifact/noise in that spot. Finetuned VAE basically has slightly different lighting contrast is pretty similar.
Finetune V2.0 settings from the script vibe-coded by Claude Sonnet 4.5 that I used:
- GPU: RTX 3060 12GB
- VAE Used: The Base Crystal VAE Merge on this page
- Dataset Total: 251 images (Overkill but Good for Variety)
- Resolution: 1024x1024
- Batch size: 2
- Epochs: 5 - Ended up choosing epoch 1 on testing, due to PC shutting off at epoch 3/5)
- Learning rate: 5e-6 (0.000005)
- Training mode: Decoder-only
- Optimizer: AdamW
- Loss: MSE reconstructionNote: Still doesn't do well in hires only for dark lighting images.
All Finetune V2.5 settings:
- GPU: RTX 3060 12GB
- VAE Used: The Base Crystal VAE Merge on this page
- Dataset Total: 55 images
- Resolution: 1024x1024
- Batch size: 1
- Epochs: 1
- Learning rate: 1e-5 (0.00001)
- Training mode: Decoder-only
- Optimizer: AdamWNote: Was gonna make a SD1.5 version but decided not to since they'd be extremely similar.
Re-Categorized due to tool due to me wondering what the Google AI overview thought was a asset.
This is what the AI Overview thought is a asset: 
This is what the AI Overview thought is a tool:

Description
Edited the saturation, brightness and contrast, slightly of V2.5.
Settings used for the saturation, brightness and contrasted: -5% brightness, 10% contrast, 2% saturation.
I recommend using this for final hires/upscale OR base image OR if you know how to use this VAE you can hires/upscale with it.
Note: Also likes to add stuff if you hires with it. According to redactedpaws.
FAQ
Comments (19)
Does this work with fp16 without upcasting?
Uhhh I made it for FP16 models which I use...all my stuff is mainly made for me to use tbf...and I use FP16 Pruned Illustrious/SDXL/NAI/Pony Models.
Edit: If it works on other stuff oh well. All I did was merge VAEs that work with FP16 CKPT of mine as the model I was testing on and Finetune the merged VAE and adjust contrast/saturation/lighting/brightness.
1 of the best vae I've used
Tensor Page with My Models: https://tensor.art/u/75677324641794
this vae is like magic. Upscaling with it don't break noisy images
Huh... Never knew that before...til now.
@Arctenox yeah, hires fix with this vae is more resistant to double navel than other vae
@kiheromasterki849 Nice.
@kiheromasterki849 我第一次听说double navel原来是VAE的问题?过去我一直把extra navel放在负面提示词,但效果不好。感谢你的发言,我也要测试一下
@wululululu 不,这不是 VAE 的问题。这主要是由于 sdxl 模型的局限性造成的。一些使用高分辨率数据集训练的组件对此的抗干扰能力更强。
@kiheromasterki849 ohh,mine is ILL,is that same?
@wululululu pony-illustrious-noob all belong to sdxl architecture
@kiheromasterki849 Yeah but this VAE was intended for Illustrious use-case, as it was only ever tested with illustrious, it was never tested with Base SDXL, NoobAI or Pony. They may be the same architecture BUT, Pony, Illustrious and NoobAI are finetunes like you can't expect to prompt the same way you do on Pony as on Illustrious as illustrious doesn't use score tags, as illustrious was trained on Danbooru. NoobAI is similar to Illustrious BUT should also alongside Danbooru tags, know E621 / A Furry Site's tags. Pony was trained on their own dataset and tags far as I am aware. But the fact is I only ever used and tested this VAE for Illustrious.
@Arctenox sdxl originally created for 1024x1024, and other finetunes were built upon that. If you try to do 1.5k resolution and above, the model will hallucinate and create artifacts like double navels. Newer models like anima or z don't have that problem
@kiheromasterki849 I know how finetunes of SDXL work lol. I use them all the time. Even got a guide. Different finetunes have different training data as well. You can't use score tags with illustrious as illustrious wont know what you are talking about so instead of score tags you use "masterpiece, best quality" and all of that. I also know Anima knows Some Booru tags and Mostly Natural Language.
I also know anima was built off of SDXL BUT because you need qwen image vae and that qwen clip it's mostly different architecture so you can't block merge Anima with Illustrious, I've tried and got a ton of errors.
But in all seriousness this VAE DOES and WILL do things in the upscale process.
@Arctenox 我知道VAE的原理是把image压缩成latent,ILL对latent进行去噪。然后VAE把完成的latent解压成新的image。所以你的意思是,你的VAE和原始的SDXL VAE,在压缩的时候不一样?latent在I2I的时候会损失原始image的信息。但是T2I会有什么改变呢?
@wululululu I am English only, so I do not know what you said. T-T
If you want to know what the VAE does in the hiresing: small saturation, clears up some noise to be clearer, minor stuff like that. That is what happens. And cause of it clearing up noise so the ckpt model in the hires-fix usually will fix it and for the upscale its usually just saturation. It doesn't do anything crazy far as I am aware from my testing otherwise it would be considered broken.
So in theory it should help prevent double-navel that is caused by blurry noise, not remove it, it will only happen if the image is a weird resolution or too big. Base SDXL only supports certain resolutions like 1024x1024 or anything that is inside the preset resolution nodes in comfyui for sdxl, illustrious and noobai support bigger image sizes usually but not always.
@Arctenox but t2i is white latent,how can VAE effect?and in i2i,I know VAE can clears a little noise,but is not that a wrong?VAE will compressed all the pixels,I thought VAE removing noise was an accident. If I need to reduce noise, I will directly lower the noise. I don't know what is the value with VAE removing noise.
@wululululu VAEs help image generation that's all there their for -> Helping Noise be Generated Properly, because without it you get white noise splotches on the Images you can try this with some models on site that have no VAE you will get faded splotchy images like it failed a red room and some models will cause some noise to not be generated blurry when it shouldve been generated not blurry VAEs that effect this only do this during t2i then hires/upscale.
VAEs that effect saturation have had saturation configurations added ontop so to speak and will always add saturation every single time it's run. Meaning do NOT run v1 of this 3 times in a hires/upscale. the saturation will be completely different from image 1 as that version already adjusts saturation and contrast heavily and far as i'm aware different vaes different minor things, just know if they don't generate an image or results are actually terrible -> do not use. If you do not know what you are doing stick with the VAE baked in a model as most illustrious/noobai/sdxl/pony models have a vae baked in. Anima however uses its own vae which I wouldn't worry about either. The only reason my VAE does what it does is because I finetuned it that way, either hate it or like not much I can do there. You can finetune VAEs similar to how you make loras aka a dataset but datasets for to finetune a VAE tend to be under 100 imgs and I did that with this VAE. Anyhow point is in the end VAEs are a must in any image generations whether it's t2i or i2i almost every one should know this heck even the creator of hyphoria knows vaes are a must and has one baked into their model.

