Qwen Image 2.1 7B INT8 and INT4 (W4A8) ConvRot Quantizations, for use in ComfyUI
Qwen-Image-2.1
Qwen-Image-2.1, A unified text-to-image generation and image editing model in the Qwen family. With just 7B parameters in its visual generation component (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility.
Highlights
Efficient Image Generation: Qwen-Image-2.1 combines strong visual performance with fast inference and a compact design, making high-quality image creation accessible across a wide range of creative workflows.
Flexible Creative Control: With support for diverse inputs, outputs, and localized edits, Qwen-Image-2.1 gives creators the flexibility to explore ideas and refine details within a unified workflow.
Four key improvements define this release:
Compact and Efficient — A lightweight architecture with mixed-granularity attention and prefix KV cache reuse delivers strong image quality at low computational cost.
Native Transparency, Unified Creation and Editing — Generate regular or transparent (RGBA) images from text, edit transparent layers, and extract subjects from photographs—all in one model.
Versatile Editing — Support up to 10 reference images, specify local edits via circles, painted annotations, or separate masks, and preserve identity for people and products.
Realistic Textures and Refined Aesthetics — Improved typography, portrait lighting, and fine details for more visually compelling results.
Description
FAQ
Comments (12)
damn Qwen has this lifeless ai vibe aesthetic in most images
Yeah, it seems like it, LoRAs are definitely needed for T2I. The edit capabilities are a bit more interesting.
@tsolful yeah loras can fix this without problem however artifacts are very difficult to fix. This image has that
As long as it's open sourced, trainable, and the base knowledge and quality is ok,
all you need is to trust the community.
@ikekph5 I trust the community 100%, I'll probably train a new fantasy realism refiner when AItoolkit supports. Non-commercial licensing tho :/ https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE
@tsolful the license is ... unfortunate ... Why just follows other qwen model like qwen3.8 with a revenue cap. Qwen image 2.1 is an old model to them anyways.
edit: GPT tells me because they don't care about the old model. But the model doesn't seem to have safety filter and is too good at nsfw so they probably want the "research or evaluation purposes only" shield for lawsuits...interesting thoughts
It's really poor in comparison of Krea 2. I hope the editing features are better. Thanks for the INT8 anyway.
the license kills this model
cant wait for the Loras, its actually good
still seeing the old halftone thing going on in image edit, so they didn't fix that in their new vae, lol
very light NSFW test showing that fine‑tuning is certainly needed, but even the bare knowledge is quite good. At the very least, there’s no censorship.
The camera angle changes are weak, but it's the best image enhancer (open source) I've seen so far. The consistency of character is very good. The screen-shot improvement is quite good.
