CivArchive
    Dataset generator Qwen2.1 - v1.2
    Preview 143643245
    Preview 143643244
    Preview 143643246

    V1.1

    Updated the dataset generation prompts to improve variation across facial expressions, lighting, camera angles, poses, and perspectives instead of keeping the face too close to the original reference expression.

    I also simplified the reference-image setup:

    • Select A for 1 reference image

    • Select B for 2 reference images

    On my RTX 5080, generating 20 images with 1 reference image of around 1 MB took roughly 10 minutes.

    Overall, V1.1 should produce a more diverse and useful dataset for LoRA training, with stronger variation in expressions, lighting, poses, and viewpoints while still preserving the subject's identity.

    Change max rows to 25 if u want 25 images and you can always change the prompts send the prompt used to chatgpt or somewhere u like to get full body shots or different styles
    -------------------------------------
    Qwen-Image 2.1 Dataset Generation Workflow
    This is an updated/adapted version of an existing dataset-generation workflow originally created by acekiube. Full credit for the original workflow structure and idea goes to them. I did not create the workflow from scratch.


    I modified the workflow to work with Qwen-Image 2.1, updated the model/encoder nodes, generation settings, resolution handling, and prompts for identity-focused dataset generation.
    The goal is to take a reference character/person and automatically generate a set of useful training images with different viewpoints, expressions, camera angles, and lighting while keeping the identity as consistent as possible.


    I’ve personally had better identity consistency around CFG 3, especially for side/profile views. The workflow is currently set to CFG 3, 25 steps, res_multistep, and a 3:4 1MP output.
    I’m also experimenting with giving Qwen additional reference views, such as a side-profile image, since this seems to help significantly when generating difficult angles.

    If you got better prompt or settings please let me know

    Models used:
    Qwen2.1 int8 convrot:
    https://huggingface.co/Comfy-Org/Qwen-Image-2.1/resolve/main/diffusion_models/qwen_image_2.1_int8_convrot.safetensors

    Text Encoder:
    https://huggingface.co/Comfy-Org/Qwen-Image-2.1/resolve/main/text_encoders/qwen3vl_8b_int8_convrot.safetensors

    Vae:
    https://huggingface.co/Comfy-Org/Qwen-Image-2.1/resolve/main/vae/qwen_image_2.1_vae_bf16.safetensors

    Description

    FAQ

    Comments (4)

    DarkarmySep 24, 2026
    CivitAI

    Really good workflow to create high quality different angles and expression with consistency. Only one thing is making trouble. I guess is the native node: ¨Qwen Image 2.1 Cache¨, first run ok takes long but normal, second run gets stuck, I guess this node has a memory leak, tried with startup .bat conditions (--disable-smart-memory --cache-none), but no go. Since text enconder gets dumped in each line of prompt it generates a loop.

    ManuelB32478Sep 24, 2026
    CivitAI

    Thanks. Works very well. It also work with a character reference sheet as input. But I was having better output with CFG=1. CFG=3 makes the output more contrasty and change the skin color.

    Jolanoff
    Author
    Sep 24, 2026

    I wanna make it create a character sheet then use it as ref to generate the dataset, could you share your wf for the sheet that's working well for you?

    ManuelB32478Sep 29, 2026· 1 reaction

    @Jolanoff sorry for the late reply. I'm using the basic qwen2.1 workflow with those prompts.

    3 panels grid:
    3-pannel grid of realistic photos, Top-left: extreme close-up of the face, 3/4 view turned to the left, gaze slightly off-camera to the left, Bottom-left: head and torso down to the waist, 3/4 view turned to the right, gaze toward the right, 3/4 view turned to the right, Right side, tall: full-body standing pose, front view, relaxed A-pose, both arms hanging down and angled slightly away from the body, remove any white grid lines or borders separating the panels — the panels should blend seamlessly into the background with no visible dividing lines, as if they were captured on a single continuous surface

    4 panels grid;

    4-panel grid of realistic photos, Top-left: extreme close-up of the face, 3/4 view turned to the left, gaze slightly off-camera to the left, Bottom-left: head and torso down to the waist, 3/4 view turned to the right, gaze toward the right, Center, tall: full-body standing pose, front view, relaxed A-pose, both arms hanging down and angled slightly away from the body so the waist and hips are fully visible between the arms and torso, Right side, tall: full-body standing pose, body turned about 70 degrees to the left in a near-profile view, head turned back toward the camera with a direct gaze, arms relaxed at the sides, weight on one leg, profile silhouette of bust, waist, hips and thighs clearly visible, remove any white grid lines or borders separating the panels — the panels should blend seamlessly into the background with no visible dividing lines, as if they were captured on a single continuous surface

    To that you add the description of the character and the environment. A full prompt will look like this:
    4-panel grid of realistic photos, Top-left: extreme close-up of the face, 3/4 view turned to the left, gaze slightly off-camera to the left, Bottom-left: head and torso down to the waist, 3/4 view turned to the right, gaze toward the right, Center, tall: full-body standing pose, front view, relaxed A-pose, both arms hanging down and angled slightly away from the body so the waist and hips are fully visible between the arms and torso, Right side, tall: full-body standing pose, body turned about 70 degrees to the left in a near-profile view, head turned back toward the camera with a direct gaze, arms relaxed at the sides, weight on one leg, profile silhouette of bust, waist, hips and thighs clearly visible, remove any white grid lines or borders separating the panels — the panels should blend seamlessly into the background with no visible dividing lines, as if they were captured on a single continuous surface, shot on 85mm f/1.4 lens, creamy background blur, a 20 year old young woman, west african, dark chocolate skin tone, balanced curvy proportions, slightly narrower waist than hips, full lips, plump volume, defined cupid's bow, smooth rounded contours, oval face, soft rounded jawline, balanced proportions, high cheekbones, monolid eyes with a smooth, flat eyelid, no visible crease, seamless transition from lid to brow bone, sleek and minimal eyelid contour, loose waves hair, black hair, neutral expression, looking at viewer, grey, cotton, t-shirt, black, cotton, leggings, converse, photo studio, soft diffused daylight, neutral grey background, photorealistic, cinematic

    Format 3:2. 1 or 2 megapixels

    Workflows
    Qwen 2

    Details

    Downloads
    544
    Platform
    CivitAI
    Platform Status
    Available
    Created
    9/24/2026
    Updated
    9/29/2026
    Deleted
    -

    Files

    datasetGeneratorQwen2_v12.json

    Mirrors