Krea 2 Turbo — QuantFunc A4W4 INT4
True 4-bit inference — A4W4 (4-bit activations × 4-bit weights).
QuantFunc's core quantized matrix multiplications use INT4 activations and INT4 weights on its INT4 inference backend. A4W4 describes the compute precision used during inference, alongside the reduced weight size and memory bandwidth demand.
QuantFunc INT4 compresses Krea-2-Turbo to about a quarter of its 16-bit size. In our visual comparisons, composition, detail, color and style stay close to the 16-bit baseline, roughly on par with FP8 and INT8 ConvRot.
Showcase
All images below were generated by Krea-2-QuantFunc-4bit.
Portrait photography Sci-fi scene
Impasto art
Watercolor illustration
Fast denoising, fast end-to-end too
High-VRAM
On an RTX 4090, same workflow and generation settings:
Stage QuantFunc INT4 FP8 Speedup Denoising 1.6s 5.0s 3.13x End-to-end 2.5s 6.5s 2.6xEnd-to-end includes text encoder, denoising and VAE. Text encoder and VAE are not part of the QuantFunc plugin's acceleration path today. Actual speed varies with resolution, steps, driver, software version and hardware.
Low-VRAM
On 8 GB / 12 GB and similar VRAM-constrained setups, 4-bit weights cut weight-bandwidth demand significantly, adding extra speedup — up to roughly 11x. Actual gains depend on VRAM capacity/bandwidth, offload behavior and generation settings.
Swap one loader, keep your workflow
- Install or update ComfyUI-QuantFunc.
- Download the r128 or r32 weights.
- Swap your model loader for the QuantFunc loader and pick the matching weight file.
Every other node, connection and generation parameter stays as-is.
Krea-2-Turbo is a distilled turbo model — use fewer sampling steps and turn CFG off (guidance 1.0).
RTX 20-series through GB300, one build covers it all
Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.
Choose a model
Variant File Size Use case r128krea2-turbo-quantfunc-int4-r128.safetensors
8.83 GB (8.22 GiB)
Quality-first, recommended
r32
krea2-turbo-quantfunc-int4-r32.safetensors
8.35 GB (7.77 GiB)
Size/VRAM-first
Loading
Weights use QuantFunc's own safetensors format — load with ComfyUI-QuantFunc or the QuantFunc inference engine, not as a drop-in Diffusers checkpoint.
Source & license
This repository distributes derived quantized weights produced from:
- Base model: krea/Krea-2-Turbo
- Pretrained base: krea/Krea-2-Raw
Weights follow the Krea 2 Community License — please read the original model's license terms before use.
Official QuantFunc links
- Hugging Face — QuantFunc
- ModelScope — QuantFunc
- Official Discord
- QuantFunc website
- ComfyUI-QuantFunc plugin and workflows
Source and weight integrity
Original QuantFunc release: QuantFunc/Krea-2-QuantFunc-4bit.
krea2-turbo-quantfunc-int4-r128.safetensors— SHA256:cf465ba39846a5664f0d92ad3c41a8303705d91cb89cbec856371cf5635066bbkrea2-turbo-quantfunc-int4-r32.safetensors— SHA256:89cddbfaeced4fafb80b0cf284bdd29aec85ca204b5a7f5f6e245331c14cc875
Showcase media is reproduced from the original QuantFunc model card. Exact seeds, prompts and rank variants are not supplied for every example. Performance figures are QuantFunc's reported measurements under the stated conditions; they are not guarantees for other workflows.
License terms
These quantized weights are a modified derivative of Krea 2 Turbo; they are not an official or endorsed Krea product. The Krea 2 Community License permits commercial use only below its USD 1 million company-wide annual revenue threshold; otherwise a separate Enterprise License is required. Redistribution must retain the agreement and NOTICE, and deployments must follow its content-filtering and acceptable-use requirements. See the full license agreement.
Description
True 4-bit inference — A4W4 (4-bit activations × 4-bit weights).
QuantFunc's core quantized matrix multiplications use INT4 activations and INT4 weights on its INT4 inference backend. A4W4 describes the compute precision used during inference, alongside the reduced weight size and memory bandwidth demand.
QuantFunc INT4 r32 quantization of Krea 2 Turbo. Smaller rank variant. Requires ComfyUI-QuantFunc or QuantFunc. Use the Turbo model schedule, reduced sampling steps and CFG 1.0.
File: krea2-turbo-quantfunc-int4-r32.safetensors
Size: 8.35 GB (7.77 GiB)
SHA256: 89cddbfaeced4fafb80b0cf284bdd29aec85ca204b5a7f5f6e245331c14cc875
Modified derivative of Krea 2 Turbo under the original Krea 2 Community License. Commercial use remains subject to the license revenue threshold.



