Quantization of Kroma v0.3 Turbo by Lodestones primarily meant for low VRAM/RAM usage, expect degradation in quality compared to INT8 Convrot only quants or BF16, saves ~64.3% in storage relative to BF16.
Spent a great deal of time working on which layers to quantize to what method (INT8 Convrot or W4A4 Convrot) and I kind of like this one.
Used a simple T2I workflow, no second stage or upscaling done as I'm too lazy for that.
Tested on:
NVIDIA GTX 1660 Super 6 GB
32 GB System RAM
Test settings:
8 steps
CFG 1
Euler / Simple
696 × 1048 resolution (0.7 on resolution selector)
Approximately 9 s/it on a GTX 1660 Super



