CivArchive
    Rebels w4a8s (ComfyUI) - All models (more coming)
    NSFW
    Preview 139344264
    Preview 139344266
    Preview 139344273
    Preview 139344269
    Preview 139344274

    Rebels w4a8 Collection

    (Check back frequently as i will update whenever i quant a new model)


    - CURRENT LIST:

    • Krea-2-Turbo

    • Z-Image-Turbo

    • Flux Klein 9B

    • Flux Klein 4B

    • Wan Animate 2 (Distilled)

    • Qwen Image 2512

    • Scail-2 (Needs Work, Not Perfect)

    • LTX-2.5 Distilled (2 versions, 1 with better audio at higher compute)

    • MiniMax-Music-3

      Four-bit weights. Eight-bit math. Native ComfyUI kernels. No custom nodes.

    W4A8 conversions of current diffusion and video models, quantized to run on ComfyUI's native asym_w4a8_int8 path — where int8 tensor cores do the work instead of a dequantize-then-fp16 fallback.

    Roughly 0.56 bytes per parameter, and it runs at that size rather than merely storing at it.


    Why W4A8 instead of GGUF

    GGUF is excellent and I ship plenty of it. But every GGUF forward pass unpacks weights back to fp16 before the matmul — the file is small, the math is not. W4A8 keeps compute in int8 end to end.

    GGUF Q4_K_MW4A8Storage~0.60 B/elem~0.56 B/elemCompute pathdequant → fp16 GEMMint8 GEMMLoaderComfyUI-GGUF nodestock Load Diffusion ModelWeight error (measured)varies by tensor~7% relL2

    The format comes from Kijai's AsymW4A8Int8Layout work in comfy-kitchen. This collection is about applying it correctly to models nobody has converted yet, and being explicit about what was verified.


    What's inside a file

    Each quantized Linear stores five pieces:

    TensorPurposeweightint4 codes, two per byteweight_s_relfp8 scale, one per group of 16weight_s_channelone scale per output channelweight_codebook16 Lloyd-Max levels, fit to the tensorcomfy_quantlayout config the loader reads

    Three ideas stacked: a ConvRot Hadamard rotation that flattens outliers so four bits go further, a codebook of non-uniform levels fit to the actual weight distribution instead of an even grid, and per-group fp8 scales preserving local dynamic range. Calibration-free — no activation dataset, so nothing in the conversion biases the model toward one kind of prompt.


    What I do differently

    Sensitive layers are never quantized. Timestep embeddings, conditioning projections, patch projections, final output layers and rotary tables stay high precision. On a few-step model the timestep embedder has only a handful of sigma values to distinguish — crushing it to four bits corrupts every step of the schedule. Every file is checked after conversion to confirm those layers really are stored at F16/F32, because quantizers do not preserve them automatically.

    Mixed formats where the kernel demands it. The fused W4A8 kernel accepts a ConvRot group of exactly 256, so any layer whose input dimension isn't divisible by 256 cannot use it. Rather than silently shipping a file that errors on load, those layers are written as int8_tensorwise — also native, no group constraint, ~1% error. Each model card states which layers took that path.

    Every file is measured. Conversion reports per-layer reconstruction error against the original bf16 weights. Anything that doesn't land where it should doesn't get uploaded.


    Requirements

    • ComfyUI 0.30.0+ with asym_w4a8_int8 in its native quant registry

    • comfy-kitchen installed (ships the kernels)

    • An NVIDIA GPU or AMD GPU.

    On startup ComfyUI prints its available formats. You want asym_w4a8_int8 in the Native ops list — under emulated it still runs, without the int8 speed advantage.


    Usage

    1. Drop the .safetensors in ComfyUI/models/diffusion_models

    2. Load it with Load Diffusion Model — the stock node, no custom loader

    3. Text encoder, VAE and sampler settings are unchanged from the base model

    Per-model notes (step counts, CFG, resolution) live in each model's card. Distilled models have fixed schedules that must be respected — base-model settings on them produce poor results regardless of quantization.


    Models

    Collection in progress. Each conversion has its own repo with exact sizes, measured error, and the list of layers kept at high precision.


    Licensing

    These are quantized derivatives. Every original license and usage restriction carries over unchanged, and each model repo states the license of its base model. Check the specific model's card before commercial use — several bases in this collection are not permissive.


    Credits

    • Kijai — the W4A8 int8-codebook layout and kernels

    • Comfy-Org / comfyanonymous — comfy-kitchen and the native quantization registry

    • city96 — ComfyUI-GGUF, which taught most of us how quantized loading works in ComfyUI

    • Original model authors — all base licenses apply

    Quantized by RealRebelAI · GitHub · X

    Description

    first set

    Comments (4)

    gou39Aug 11, 2026
    CivitAI

    Kreamagine_LoRA どこにあるのですか

    realrebelai
    Author
    Aug 11, 2026

    private collection

    vladulidloAug 11, 2026
    CivitAI

    Would you please add LTX 2.3 DEV W4A8?

    realrebelai
    Author
    Aug 11, 2026· 1 reaction

    2.5 is coming today

    UNet
    Other

    Details

    Downloads
    42
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/10/2026
    Updated
    8/30/2026
    Deleted
    -

    Files