Quantization of Kroma v0.3 Turbo by Lodestones primarily meant for low VRAM/RAM usage, expect degradation in quality compared to INT8 Convrot only quants or BF16, saves ~64.3% in storage relative to BF16.
Spent a great deal of time working on which layers to quantize to what method (INT8 Convrot or W4A4 Convrot) and I kind of like this one.
Used a simple T2I workflow, no second stage or upscaling done as I'm too lazy for that.
Tested on:
NVIDIA GTX 1660 Super 6 GB
32 GB System RAM
Test settings:
8 steps
CFG 1
Euler / Simple
696 × 1048 resolution (0.7 on resolution selector)
Approximately 9 s/it on a GTX 1660 Super
Description
FAQ
Comments (9)
There would be int8 12gb model?
You can find them in huggingface, though int8 convrot only is about 14 gb.
The model follows instructions quite well. However, the images often turn out overly glossy or "soapy" And when I specify "a wild beauty gazing into the distance," the resulting expressions sometimes wouldn't look out of place in a horror movie—but they certainly aren't "beautiful."
Yeah, that's the downside with these heavily quantized models. Quality definitely takes quite a hit, I spent a good amount of time trying to reduce that where it matters the most (facial features, hands etc) but at that size, it is what it is I guess.
And I noted the glossy part too, I can't be sure of this but it's a bit similar with the int8 convrot only model as well so I am assuming it's a model thing. Though you can write certain keywords or change the prompting style to nudge it in a different direction.
I could upload int8 convrot or a mixed of around 10 gb if you would like to test a higher quality one.
@BakaPotatoLord Ok, i would test it.
damn, could you do the same mixed quant for just the standard krea2 turbo? under 9gb is so nice
If you are fine with the quality loss, sure can do
@BakaPotatoLord that would be cool, yea. Thanks!





