Built with Qwen. A few-step distilled version of Qwen/Qwen-Image-2.1 by Viggle. Text-to-image and instruction-driven editing with 1–3 reference images in 6 steps instead of 40, with no classifier-free guidance.
About 5× faster than the 40-step base model end to end, and very competitive with it in quality: on the official Qwen examples the two are hard to tell apart on most prompts.
Original Model card: https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo
Description
6 steps, a different balance. Against v0.2.1: less grain and cleaner surfaces, fine texture a little softer; diversity (still close to the base model's) and small-text accuracy about the same.
New 9-step mode: 7 turbo steps, then the LoRA is switched off and the base model finishes the last two. Finer detail, and small text comes out right more often (not always). It takes about 1.4–1.5× as long as 6 steps (still about 3.5× faster than the base model). It works in diffusers and the demo Space only (9 steps).
We think 6 steps is close to its capacity. Since v0.2.1, every gain we found at 6 steps cost something elsewhere: sharper came with more grain, less grain came with a softer look. Beyond this point, quality most likely has to be paid for with steps, which is what the 9-step mode does.


