CivArchive
    MiniMax-H3 Ref2VA (4-Bit INT4 Safetensors) - Ref2VA-DF-Turbo-W4A8
    Preview 140536113

    # ⚡ MiniMax-H3 Ref2VA (Safetensors Quantized Checkpoints)

    This repository provides optimized Safetensors quantizations of MiniMax-H3-Ref2VA (Reference-to-Video & Audio), engineered specifically for consumer GPUs (16 GB VRAM, RTX 4080 / RTX 4090 / RTX 4090 Mobile / RTX 3090).

    Choose between two purpose-built variants depending on your workflow and VRAM budget:

    ---

    ## 🌟 Variant 1: Digital Forge Turbo W4A8 (All-In-One 4-Step) — This is technically coming soon, this is currently an offload compression version, v2 will be the full model. To be frank, an issue with the audio was discovered. Video quality was there, but audio took a hit. Have to make some changes to the quant process to for the audio. So, coming v2, coming soon. For now the int4 is usable, the other is just a compression offload model.


    The flagship mixed-precision release designed by Digital Forge (DF). It eliminates common INT4 artifacts while mathematically fusing the official LightX2V 4-Step Turbo adapter directly into the model weights.

    ### 🛡️ Why Choose DF-Turbo-W4A8?

    🚀 *Pre-Baked 4-Step Turbo**: The 4-step Turbo LoRA is fused before quantization. No external LoRA loading required—saves 1.28 GB runtime VRAM and eliminates 624 GEMMs per step!

    🛡️ *Zero "Zombie Sway" (Pristine BF16 AdaLN)**: All 335 time-modulation layers, LayerNorms, and 1,025 calibration grid points remain in unquantized BF16, ensuring continuous temporal velocity ($\vec{v}$) and smooth, natural character motion.

    🎯 *Zero Reference Feature Drift (INT8 Cross-Attention)**: Uses INT8 per-channel quantization across all cross-attention to_k and to_v layers to preserve fine facial features, eye contact, and clothing textures without outlier clipping.

    💾 *16 GB VRAM Resident (12.62 GiB)**: Fits 100% resident inside 16 GB GPUs with zero PCIe swapping.

    ### ⚙️ Recommended Settings (DF-Turbo-W4A8)

    * Sampling Steps: 4

    * CFG Guidance: 1.0 (Distilled trajectory)

    * Video Shift: 12.0

    * Audio Shift: 3.0 (32 kHz synchronized audio)

    * Frame Count: 25 to 124 frames

    ---

    ## 🧪 Variant 2: Pure INT4 (Ultra-Low VRAM Experimental)

    A pure 4-bit uniform quantization designed for minimal memory consumption.

    ### 🔬 Model Details & Caveats

    * VRAM Footprint: *9.80 GiB** (Ultra-compact, maximum memory headroom).

    * Quantization Scheme: Symmetric 4-Bit Linear FastInt4Linear) with Group-128 scaling.

    * Adapter Support: Requires loading the external LightX2V 4-step Turbo LoRA minimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors) if 4-step generation is desired.

    ⚠️ *Experimental Note**: In extended runs (e.g. 10-second / 124+ frame sequences), pure INT4 quantization on time-modulation layers may exhibit slight identity drift or rhythmic "zombie sway" artifacts. Use DF-Turbo-W4A8 if artifact-free motion is required.

    ---

    ## ⚡ Performance & Engine Compatibility (Both Models)

    * Zero-Copy Loading: Loads via OS virtual memory mapping mmap) in ~0.03 to 0.05 seconds.

    * TeaCache Compatible: Full native support for Block-Level TeaCache (~50% compute step bypass).

    * Attention Acceleration: Supports SageAttention 2.0 Patched (Fast FP8 with FP32 accumulator), FlashAttention-2, and native PyTorch SDPA.

    * VAE Tiling: Recommended VAE tile size of 512 for efficient 3D temporal decoding.

    ---

    ## 📜 License & Attribution

    Base model licensed under the *MiniMax H3 Community License Agreement**, Copyright © 2026 MiniMax.

    Powered by *MiniMax H3** & Digital Forge.

    Description

    Checkpoint
    Other

    Details

    Downloads
    50
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/23/2026
    Updated
    8/24/2026
    Deleted
    -

    Files

    minimaxH3Ref2va4Bit_ref2vaDFTurboW4A8.safetensors