About these variants
Four precision levels are available for Qwen-Image-2.1, trading VRAM and speed against generation quality:
- BF16 — full-precision reference (~14 GB). Highest quality, largest footprint.
- INT8 (W8A8) — 8-bit weights + 8-bit activations (~7 GB). Near-lossless quality, ~2× smaller than BF16, runs on INT8 tensor cores (RTX 30-series and up).
- INT6 (W6A8) — 6-bit weights + 8-bit activations (~5.5 GB). Middle-ground: smaller than INT8 with only a small quality trade-off.
- INT4 (W4A8) — 4-bit weights + 8-bit activations (~4 GB). Smallest footprint; uses ConvRot with a per-tensor codebook that decodes to INT8 for compute, so it runs on the same INT8 hardware as W8A8. Larger quality trade-off than INT6.
Qwen-Image-2.1 7B — INT4 (W4A8) & INT6 (W6A8) ConvRot for ComfyUI
INT4 and INT6 quantized weights of Qwen-Image-2.1 for fast, low-VRAM inference in ComfyUI. This is a modified (quantized) version of the Qwen-Image-2.1 model. It is not an official Qwen release and is not endorsed by the Qwen team.
About Qwen-Image-2.1
A unified text-to-image generation and image editing model in the Qwen family. With just 7B parameters in its visual generation component (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility.
Highlights
- Efficient Image Generation — combines strong visual performance with fast inference and a compact design, making high-quality image creation accessible across a wide range of creative workflows.
- Flexible Creative Control — supports diverse inputs, outputs, and localized edits, giving creators the flexibility to explore ideas and refine details within a unified workflow.
Key improvements in 2.1
- Compact and Efficient — lightweight architecture with mixed-granularity attention and prefix KV cache reuse delivers strong image quality at low computational cost.
- Native Transparency, Unified Creation and Editing — generate regular or transparent (RGBA) images from text, edit transparent layers, and extract subjects from photographs — all in one model.
- Versatile Editing — support up to 10 reference images, specify local edits via circles, painted annotations, or separate masks, and preserve identity for people and products.
- Realistic Textures and Refined Aesthetics — improved typography, portrait lighting, and fine details for more visually compelling results.
License
Qwen-Image-2.1 is licensed under the [Qwen Research License] These files are a quantized derivative and are distributed under the same license — see the original repository for full terms, permitted uses, and any commercial-use restrictions.



