INT8 Mix Quantization of One Obsession v24 by maxfeifei8, I got really curious about how INT8 quantization would fare with Illustrious. Turns out pretty good.
70 layers in INT8 and rest in INT8 Convrot for the diffusion model.
v2.0
Text Encoders are now quantized with a mix of INT8 and INT8 Convrot, total checkpoint size is now down to 3.92 GB
A patch is required to run the model as ComfyUI does not support quantization methods for the SDXL Clip Loaders. YOU WILL GET GARBAGE OUTPUT WITHOUT IT.
Same prompt and seed from v1.0 so you can compare the images in terms of quality
v1.0
Text Encoders are untouched
Checkpoint runs natively on ComfyUI
Clone the below repo into your ComfyUI's custom_nodes folder. No need to add any nodes, it initializes and patches the CLIP Loader at runtime, please let me know if you encounter any issues.
https://github.com/bakapotatolord/ComfyUI-PotatoForge-Encoder-Quant
Sample images once again just the first stage (T2V), used prompts from the original model.
Tested on:
NVIDIA GTX 1660 Super 6 GB
32 GB System RAM
Test settings:
25 steps
CFG 4.5
Euler Ancestral / Normall
696 × 1048 resolution
Approximately 2.5 s/it on a GTX 1660 Super, about 45% faster than the usual model
Description
Quantized the text encoders (CLIP L and Clip G), bringing down the size to 3.92 GB
As ComfyUI does not support quantization methods for the SDXL Clip Loader, a patch is required, just clone the below link into custom_nodes/
https://github.com/bakapotatolord/ComfyUI-PotatoForge-Encoder-Quant
No node or anything required, patches during initialization
FAQ
Comments (7)
Illustrious's strength is its speed, so it must have become even faster.
Oh yep, it's twice as fast on my GTX 1660 Super. v1.0 is more than enough for just speed up, v0.2 is more of a reducing storage kind of thing.
Performance improved even further with the even lower specifications of an NVIDIA GTX 2600 Super (6 GB) and 8 GB of memory.
I look forward to working with you again.
Good to know!
Btw, do you mean RTX 2060?
Thanks man, if i did not seen you do this i would not even bother searching for methods to make Illu an int8conv (as i would assume its not possible).
Althou it ended up not being as space saving as i hoped for 6.40 to 5.00 and being faster by slightest margin (10% give or take) it is still smaller and still faster so perfect in multi-step workflows, right? :)
It's 4.43 GB for just diffusion int8 convrot and 3.92 GB for diffusion + encoder quantized, decent savings if you have a lot of models, it adds up.
But yeah, only someone using these models for their work in some professional setting can tell how bad the quality has dropped.
I think I can actually try a INT6 - INT8 mixed for Illu and see how that goes, it should reduce storage even further but the quality will surely take a hit lol
@BakaPotatoLord yeah :) i wanted to get my hand on int8 format as its simply faster while quality drop on int8conv shouldn't be more than 4% at worse and in reality of course i tested and quality drop is more like 1% when used as hires fix.
Int8Conv is truly amazing format for diffusion models ദ്ദി(˵ •̀ ᴗ - ˵ )



