CivArchive
    Krea 2 Turbo — QuantFunc A4W4 INT4 - A4W4 INT4 r128
    Preview 144379935
    Preview 144379936
    Preview 144379937
    Preview 144379938

    Krea 2 Turbo — QuantFunc A4W4 INT4

    True 4-bit inference — A4W4 (4-bit activations × 4-bit weights).

    QuantFunc's core quantized matrix multiplications use INT4 activations and INT4 weights on its INT4 inference backend. A4W4 describes the compute precision used during inference, alongside the reduced weight size and memory bandwidth demand.

    QuantFunc INT4 compresses Krea-2-Turbo to about a quarter of its 16-bit size. In our visual comparisons, composition, detail, color and style stay close to the 16-bit baseline, roughly on par with FP8 and INT8 ConvRot.

    Showcase

    All images below were generated by Krea-2-QuantFunc-4bit.

    Portrait photography Sci-fi scene Portrait Sci-fi Impasto art Watercolor illustration Impasto Watercolor

    Fast denoising, fast end-to-end too

    High-VRAM

    On an RTX 4090, same workflow and generation settings:

    Stage QuantFunc INT4 FP8 Speedup Denoising 1.6s 5.0s 3.13x End-to-end 2.5s 6.5s 2.6x

    End-to-end includes text encoder, denoising and VAE. Text encoder and VAE are not part of the QuantFunc plugin's acceleration path today. Actual speed varies with resolution, steps, driver, software version and hardware.

    Low-VRAM

    On 8 GB / 12 GB and similar VRAM-constrained setups, 4-bit weights cut weight-bandwidth demand significantly, adding extra speedup — up to roughly 11x. Actual gains depend on VRAM capacity/bandwidth, offload behavior and generation settings.

    Swap one loader, keep your workflow

    1. Install or update ComfyUI-QuantFunc.
    2. Download the r128 or r32 weights.
    3. Swap your model loader for the QuantFunc loader and pick the matching weight file.

    Every other node, connection and generation parameter stays as-is.

    Krea-2-Turbo is a distilled turbo model — use fewer sampling steps and turn CFG off (guidance 1.0).

    RTX 20-series through GB300, one build covers it all

    Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.

    Choose a model

    Variant File Size Use case r128 krea2-turbo-quantfunc-int4-r128.safetensors 8.83 GB (8.22 GiB) Quality-first, recommended r32 krea2-turbo-quantfunc-int4-r32.safetensors 8.35 GB (7.77 GiB) Size/VRAM-first

    Loading

    Weights use QuantFunc's own safetensors format — load with ComfyUI-QuantFunc or the QuantFunc inference engine, not as a drop-in Diffusers checkpoint.

    Source & license

    This repository distributes derived quantized weights produced from:

    Weights follow the Krea 2 Community License — please read the original model's license terms before use.

    Official QuantFunc links

    Source and weight integrity

    Original QuantFunc release: QuantFunc/Krea-2-QuantFunc-4bit.

    • krea2-turbo-quantfunc-int4-r128.safetensors — SHA256: cf465ba39846a5664f0d92ad3c41a8303705d91cb89cbec856371cf5635066bb
    • krea2-turbo-quantfunc-int4-r32.safetensors — SHA256: 89cddbfaeced4fafb80b0cf284bdd29aec85ca204b5a7f5f6e245331c14cc875

    Showcase media is reproduced from the original QuantFunc model card. Exact seeds, prompts and rank variants are not supplied for every example. Performance figures are QuantFunc's reported measurements under the stated conditions; they are not guarantees for other workflows.

    License terms

    These quantized weights are a modified derivative of Krea 2 Turbo; they are not an official or endorsed Krea product. The Krea 2 Community License permits commercial use only below its USD 1 million company-wide annual revenue threshold; otherwise a separate Enterprise License is required. Redistribution must retain the agreement and NOTICE, and deployments must follow its content-filtering and acceptable-use requirements. See the full license agreement.

    Description

    True 4-bit inference — A4W4 (4-bit activations × 4-bit weights).

    QuantFunc's core quantized matrix multiplications use INT4 activations and INT4 weights on its INT4 inference backend. A4W4 describes the compute precision used during inference, alongside the reduced weight size and memory bandwidth demand.

    QuantFunc INT4 r128 quantization of Krea 2 Turbo. Quality-first variant. Requires ComfyUI-QuantFunc or QuantFunc. Use the Turbo model schedule, reduced sampling steps and CFG 1.0.

    File: krea2-turbo-quantfunc-int4-r128.safetensors
    Size: 8.83 GB (8.22 GiB)
    SHA256: cf465ba39846a5664f0d92ad3c41a8303705d91cb89cbec856371cf5635066bb

    Modified derivative of Krea 2 Turbo under the original Krea 2 Community License. Commercial use remains subject to the license revenue threshold.

    FAQ

    Checkpoint
    Krea 2

    Details

    Downloads
    85
    Platform
    CivitAI
    Platform Status
    Available
    Created
    10/1/2026
    Updated
    10/5/2026
    Deleted
    -

    Files

    krea2-turbo-quantfunc-int4-r128.safetensors