Changelog
V3.0 [FINAL FOR QWEN IMAGE 2.1]
This will be my final update for the Qwen Image 2.1 lineup. There is little reason to add or tweak further, so I am officially handing the torch over to the community.
The major concept shift in this release is native support for static prompts. I've seen many comments asking: "Why use a PE (Prompt Enhancer) when you can just send a static prompt?" To an extent, that's completely valid - if you're not a writer or artist and simply want to quickly animate your character, static prompts are the way to go.
However, I kept the Auto-PE mode for users who still need granular control over character traits, outfit consistency, and fine details.
PE Model Revamp
A major change in this version is the overhaul of the PE pipeline itself. For many, GGUF proved to be too slow, while running Qwen 3.5 9B natively was too VRAM-heavy. I decided to take a different approach and switch to Qwen 3.5 4B Safetensors.
Why this model?
Tests showed that Qwen 3 VL (4B/8B) struggled with complex instruction following, frequently repeating the system prompt or outputting low-quality sheets.
Switching to Qwen 3.5 4B eliminated the repetition issue, though initial output lacked detail.
I fixed this by streamlining the system prompt. Enabling thinking mode produced noticeably better results: the model reasons through scene composition step-by-step (though it takes a bit more generation time).
Ultimately, prompt generation now takes between 30 seconds and 2 minutes (for int8), with much higher reproducibility and fewer project dependencies.
Two Workflows in One Bundle
I combined V1 and V2 into a single unified package containing two specialized workflows:
..._Production- Use this if you want a polished, highly detailed, and aesthetically rich Character Sheet...._Simple- Use this if you want a streamlined, production-ready Character Sheet built strictly for video animation.
Usage Recommendations
PE Model Selection: You can use other PE models (local or cloud-based). However, for local deployment, I highly recommend Qwen 3.5 4B / 9B depending on your available VRAM. On an 8GB VRAM setup using Qwen 3.5 4B int8, I get around 30 tokens/sec.
Static Prompts: If you don't need manual control over character backstory or precise details, stick to static prompts to speed up generation significantly.
Mode PE Static Prompt Production + - Simple + +
Performance Tip: If PE token generation speed drops unexpectedly, try re-queuing the generation. This appears to be a caching issue, and restarting restores normal speed.
Full Changelog (V3.0)
Dropped GGUF support in favor of native Safetensors.
Rewrote system prompt to optimize instruction adherence for smaller LLM sizes.
Merged V1 (Production) and V2 (Simple) workflows into a single release package.
Resource Links
Qwen 3.5 4B int8: HuggingFace | Qwen 3.5 4B bf16: HuggingFace
Qwen 3.5 9B int8: HuggingFace | Qwen 3.5 9B bf16: HuggingFace
V2.0
In version 2.0, based on community feedback, I redesigned the workflow to be much more task-focused: minimal unnecessary typography, maximum focus on key character elements, body structure, and facial details. This layout makes it drastically easier for video generation models to read character details and maintain consistency.
How It Works
The PE (Prompt Enhancer) generates 2 to 3 main panels:
Facial Expressions
Poses
Equipment & Accessories (optional)
This amount of information is ideal for video models to produce coherent, high-quality results.
Layout Rules & Features:
Faceless Entities: If the object or character does not have a human face (e.g., mask, robot, inanimate object), only a single neutral close-up panel is generated.
Expression Variety: The set of emotional expressions adapts to the character's backstory and description, but a neutral expression always comes first.
Gear & Accessories: If accessories or gear are present, a separate 3rd panel displays them in detail, isolated from the main character.
character_descriptionParameter: Helps the PE better understand character traits and posture. For instance, providing a description like "Ayaka is a slender, fragile girl. She is mostly shy, but occasionally gives a subtle smile" allows the PE to choose matching expressions and poses. This parameter is optional (the PE can infer details strictly from the image), but manually specifying details yields much higher accuracy.
Important Considerations:
Input Image: The workflow performs best when the reference image shows a full-body character facing forward. The PE is intentionally constrained from hallucinating missing body parts unless explicitly specified in the
character_description. This prevents unwanted inconsistencies like altered height, age, or outfit details.Two Workflow Variants: They differ only in how the PE prompt is generated:
For GPUs with <12GB VRAM: Use the standard version (without the
_native_pesuffix).The official
Generate Textnode is currently unoptimized for Qwen 3.5 9B, which can cause prompt generation on lower-end hardware to exceed 30 minutes. The developers are aware of this issue and working on performance fixes.
Summary of Changes (v2.0):
Complete Redesign: Shifted focus from purely stylistic/artistic design sheets to a functional, production-ready Character Design Sheet optimized for video models.
Updated Text Field: Replaced
entity_name(name only) withcharacter_description(full personality & detail specifications).Native PE Support: Added a secondary workflow utilizing the official PE execution method.
Usability: Added clear explanatory comments inside the ComfyUI node graphs.
About this Workflow
This workflow allows you to generate detailed Character Design Sheets from a single reference image. These sheets can later be used as reference frameworks for video generation models like Seedance 2.5 or MiniMax H3.
In my experience, Qwen Image 2.1 is the first open-source image editing model capable of natively generating clean, highly detailed character design sheets right out of the box.
Recommended Settings
CLIP Encoder: Use FP32 / FP16 / BF16 precision for maximum image quality and prompt adherence. Tests show that INT8 and other lower-bit quantizations introduce noticeable artifacts and reduce image clarity. Don't skimp on SSD space — install qwen3vl_8b_bf16.safetensors.
Sampler & Scheduler: Use res_2m + beta to reduce artifacts and improve detail. While I haven't run extensive benchmarks on every combination, this pairing yielded the best results.
Resolution (Megapixels):
3.4 MP — Sweet spot for speed and clarity.
6.0 MP — Superior quality and detail, though generation time increases significantly.
How to Use
Install custom nodes via ComfyUI Manager:
Download Qwen PE I2I or Qwen 3.5 4B/9B weights and place them into
ComfyUI/models/text_encoders/.Run generation: Load your reference image, select the path to the PE model, and start generating.
Useful Resources
Official Qwen Image 2.1 Weight for ComfyUI: HuggingFace
Qwen PE I2I: HuggingFace
Qwen 3.5 4B int8: HuggingFace | Qwen 3.5 4B bf16: HuggingFace
Qwen 3.5 9B int8: HuggingFace | Qwen 3.5 9B bf16: HuggingFace
Description
FAQ
Comments (15)
Considering a lot of people will use these references to feed back into the model, is there any benefit to the macro views and text labels?
Ok. Gave it a shot. It's a cool workflow but in terms of utility? You're better off a simpler character sheet. No text, white background, four views -Close up-FullbodySide-FullbodyFront-FullBodyBehind.
@AirbagGuy You’re right about that. A sheet like that might be suitable only for the artists working on the character. For video models, I’d recommend using a simpler layout. I’m currently working on a new version of the workflow that incorporates these adjustments.
@NeuroContent I think that’d be wise. I might take a crack at it this weekend myself. Maybe try a quad run sort of thing. I imagine you could use one llm run and split it up.
@AirbagGuy That's a pretty smart solution. I'd be happy if you could make it happen! By the way, I have already developed the second version of the workflow taking into account the requirements, so all I have to do is add comments, a description, and post it here.
@NeuroContent Beat me to it! https://civitai.red/models/2966369/qwen-21-character-sheet-builder-no-llm
Nice, the image design, positions, etc. are about 95% identical to the design generated by my character sheet LoRa. I intend to train one using the dataset I have here.
Character Sheets - Krea2 - Dynamic | Krea 2 LoRA | Civitai
Yes, I had to rework your prompt so that it better aligns with the official Qwen Image system prompt. Thank you for the work you’ve done! You’ve helped so many people with your projects!
First of all, thanks for sharing ❤️
The results are pretty impressive for sure!
But... it took about 16 minutes long to generate the TOKENS, I followed the same recommended models you mentioned, not sure what's wrong with "Character Sheet Prompt Maker" but it's extremely slow.
After that the Generation itself took about 4 Minutes, so yeah total of 20 Minutes for a single RUN... I guess changing the pe_i2i model to something else than what you recommended will reduce the accuracy and not sure if it even improve much of the token generation, I'm not trying because I don't want to waste extra time so maybe you'll have a good advice for me.
I'm on RTX 5090 32GB and it's the slowest workflow I've ever gave a chance, usually a single generate with Qwen Image 2.1 on 2K is about 34-37 seconds, but this is insane, 16 minutes... until the specific node finished it's Generating Tokens.
Consider I any advice for the none-GGUF version which is what I use, it's not even using much of the VRAM about 12.9GB VRAM.
Any advice? 🙏
You used the "Character_Sheet_Simplify_native_pe" workflow provided in the same archive as the GGUF version. For your computer configuration, this should speed up token generation significantly. If you still want to use GGUF PE, I recommend checking the gpu_layers parameter in the Character Sheet Prompt Maker node - it should be set to "-1."
@NeuroContent This. Use the native_pe version. I also have a 5090 and generation time was 5 minutes at 6mp
@juiceman868 I'm glad everything worked out for you. 😀
@juiceman868 so you're using the GGUF one?
@Virtual I'm not using a GGUF. I just loaded the V2.0 of the workflow, and used its default settings. The V2.0 doesn't need the GGUF, that was needed by the V1.0 version I believe.
Edit: I just realized I do have the GGUF installed in the location outlined in the notes, I must have done it for V1.0. I'm not sure if its being used. But I do have the GGUF ComfyUI/models/LLM/GGUF/
If you switch to qwen3 vl 8B for generating the prompt you can get good results about 90% of the time in around a minute on a 5090. Otherwise it should be a little under 5 minutes using a PE clip model for prompt generation.



