Experimental INT4 GGUF of Krea 2 Turbo offering decent visual output quality, 42% reduction in model size (7GB) and 35-59% performance improvement over INT8 safetensors.
IMPORTANT!
Please update the ComfyUI GGUF loader nodes (this fork: https://registry.comfy.org/publishers/molbal/nodes/comfyui-gguf-reboot) to load this properly. You will need version v28.09.04 or newer and use the Unet loader (GGUF, Dynamic VRAM) node to load it.
INT4 GGUF speed on my Windows and Linux machines, compared to INT8 ConvRot safetensors:


(Numbers, hardware, and testing setup in the article)
More information in the information page here: https://molbal.github.io/gguf/ecosystem/quant-formats.html#q4-cr-int4-convrot
Description
Comments (4)
Hello,
does this require specific hardware? Using the Unet loader I get the error "ValueError: Unexpected architecture type in GGUF file: 'krea2'"
Im on an older 1080TI, so that might be an issue?
Hi, I have not tried it with a 1080Ti, but make sure you use the latest comfyui-gguf nodes (not city96, but this one: https://registry.comfy.org/publishers/molbal/nodes/comfyui-gguf-reboot )
Should we be using "Unet Loader (GGUF/Advanced)" or "Unet Loader (GGUF, Dynamic VRAM)"?
With the first node I see a slight improvement in speed over int8 (1.28it/s over 1.02it/s) but there is a long delay each time as the model loads (even though it says "loaded completely").
With the second node the model stays loaded, but the speed is the same as int8.
Use the dynamic VRAM one. That loading should be one time only and only if you use LoRAs. So if you use the same set of LoRAs and tweak the prompt then it should not do the loading.
That improvement of ~25-30% is the benefit that should be it.
May I ask how many LoRAs you use and what is your GPU?














