CivArchive
    MiniMax H3 — QuantFunc A4W4 INT4 - FL2VA · A4W4 INT4 r128 · 4 steps
    NSFW
    Preview 144379906
    Preview 144379907
    Preview 144379908
    Preview 144379909

    MiniMax H3 — QuantFunc A4W4 INT4

    True 4-bit inference — A4W4 (4-bit activations × 4-bit weights).

    QuantFunc's core quantized matrix multiplications use INT4 activations and INT4 weights on its INT4 inference backend. A4W4 describes the compute precision used during inference, alongside the reduced weight size and memory bandwidth demand.

    QuantFunc A4W4 uses 4-bit weights and 4-bit activations for MiniMax H3's core quantized matrix computations. Character detail, style fidelity and fast motion stay clear and coherent, while H3's native video+audio generation is fully preserved.

    In our internal FL2VA evaluation, QuantFunc INT4 vs the BF16 baseline (same prompt, same seed) measures ~23.7 dB PSNR.

    Showcase

    All clips below were generated by MiniMax-H3-QuantFunc-4bit — click a player to watch.

    Live-action performance Animated chase MiniMax H3 generated video preview
    Watch generated video with audio MiniMax H3 generated video preview
    Watch generated video with audio High-speed car chase Stylized character MiniMax H3 generated video preview
    Watch generated video with audio MiniMax H3 generated video preview
    Watch generated video with audio

    Poster frames are the first frame of each clip. Showcase spec: 896 × 1184, 124 frames, 24 FPS, ~5s.

    Same 124 frames, up to 3.19x FP8's per-step speed

    On an RTX 4090, 768 × 768, 5s, 124 frames:

    Backend Per-step time Relative speed QuantFunc INT4 3.2s — INT8 ConvRot 8.5s 2.66x faster FP8 10.2s 3.19x faster

    Per-step core-model inference time only — excludes text encoder, VAE, audio processing and video saving. Actual speed varies with resolution, frame count, driver and software version.

    Swap one loader, keep the rest of your workflow

    1. Install or update ComfyUI-QuantFunc.
    2. Download the FL2VA 4-step or Ref2VA 8-step weights.
    3. Swap your model loader for the QuantFunc loader and pick the matching weight file.

    Prompts, reference images/videos, audio assets and the rest of your workflow nodes stay unchanged.

    Both weight sets already have the acceleration LoRA, Token Refiner, INT4 Refiner and INT8 Conv Sidecar fused in — no extra components to attach.

    RTX 20-series through GB300, one build covers it all

    Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.

    4-bit weights significantly cut the weight-bandwidth cost of loading and inference, making MiniMax H3 much easier to run on consumer GPUs. Actual VRAM needs depend on resolution, frame count, reference-asset count and the rest of your workflow.

    Choose a model

    Variant File Size Recommended steps Use case FL2VA minimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors 12.37 GB (11.52 GiB) 4 steps Text-to-audio/video, first-frame, last-frame, first+last-frame control Ref2VA minimax_h3_ref2va_8steps_quantfunc_int4_r128.safetensors 12.37 GB (11.52 GiB) 8 steps Multimodal reference generation from images, video and audio

    FL2VA supports zero, one or two input images; Ref2VA targets more complex multimodal reference scenarios. See the official MiniMax H3 repo for details on both modes.

    Use the matching scheduler for the model. Both weight sets already carry the acceleration LoRA fused in — no need to load it separately.

    Loading

    Weights use QuantFunc's own sealed safetensors format (quantization parameters and metadata are sealed).

    Load with ComfyUI-QuantFunc or the QuantFunc inference engine — this is not a drop-in Diffusers checkpoint.

    If the model isn't recognized or fails to load, update ComfyUI-QuantFunc to the latest version and restart ComfyUI.

    Technical details

    • INT4 weights × INT4 activations (W4A4)
    • SVDQuant rank 128, group size 64
    • INT4 Token Refiner + INT4 Refiner
    • INT8 Conv Sidecar
    • FL2VA measured max relative output error from fp16 adaLN folding: ~7.58e-4
    • Minimum GPU architecture: NVIDIA SM75

    Source & license

    This repository distributes derived quantized weights produced from MiniMax H3.

    MiniMax H3 and its derivative weights follow the MiniMax H3 Community License Agreement. QuantFunc's quantization tooling and implementation follow their own respective licenses. Please read and comply with the original model's license terms before use.

    Official QuantFunc links

    License terms

    The MiniMax H3 Community License covers these derived weights. Its standard grant excludes the European Union, United Kingdom, Republic of Korea and United States. Separate authorization is required where the community grant does not apply. Commercial products/services above the license's revenue threshold also require prior written authorization. Redistribution must include the license agreement and its required NOTICE. Full license.

    Source and weight integrity

    Original QuantFunc release: QuantFunc/Minimax-H3-Quantfunc-4bit.

    • minimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors — SHA256: fa9526b0891b455e63e293efc52331a82bc69cd3a3787625b49ccbd8145c5544
    • minimax_h3_ref2va_8steps_quantfunc_int4_r128.safetensors — SHA256: 053d6d4e7d85ee18d0f0784eb1e28774ed0fef8a1f3175a8dd451eff2254b258

    Showcase media is reproduced from the original QuantFunc model card. Exact seeds, prompts and rank variants are not supplied for every example. Performance figures are QuantFunc's reported measurements under the stated conditions; they are not guarantees for other workflows.

    Description

    True 4-bit inference — A4W4 (4-bit activations × 4-bit weights).

    QuantFunc's core quantized matrix multiplications use INT4 activations and INT4 weights on its INT4 inference backend. A4W4 describes the compute precision used during inference, alongside the reduced weight size and memory bandwidth demand.

    FL2VA: text-to-audio/video and zero, one or two input frames; use the matching 4-step scheduler.

    Requires ComfyUI-QuantFunc or QuantFunc. Acceleration LoRA, Token Refiner, INT4 Refiner and INT8 Conv Sidecar are already fused. No separate acceleration LoRA is required.

    File: minimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors
    Size: 12.37 GB (11.52 GiB)
    SHA256: fa9526b0891b455e63e293efc52331a82bc69cd3a3787625b49ccbd8145c5544

    Derived from MiniMax H3; retain its applicable license and attribution requirements.

    FAQ

    Comments (3)

    FrogOnStiltsOct 2, 2026
    CivitAI

    Hey 🐸
    Would you mind explaining why it needs a custom loader of yours? Thank you.

    QuantFunc
    Author
    Oct 2, 2026

    Since this is a new algorithm architecture that ComfyUI doesn’t support, I decided to build it myself. I plan to offer paid SDKs for languages such as Python and Java, while keeping the ComfyUI model loader free to use. This will help me continue improving the workflow and supporting the community.

    artbaseOct 2, 2026· 1 reaction
    CivitAI

    I kind a hate it where this is going. Sure, your algorithm is neat, but why does it even have an API key that checks online every time it starts? Right, this is a Beta test and later you want to monetize it. Why tho? The only guys that deserve any kind of monetary compensation are the guys who made ComfyUI and the guys who made Minimax H3. Sooner or later your method/algorithm will be outdated and the new ones will be (most likely) usable for free - locally.

    Checkpoint
    MiniMax H3

    Details

    Downloads
    151
    Platform
    CivitAI
    Platform Status
    Available
    Created
    10/1/2026
    Updated
    10/5/2026
    Deleted
    -

    Files

    minimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors