❤️ If you enjoy my work and would like to support me, consider buying me a
☕ Coffee.
This release provides a ComfyUI-compatible safetensors package of Huihui-Qwen3-VL-4B-Instruct-abliterated, converted into the format expected by the Krea 2 text encoder loader.
The purpose of this upload is simply to make the model easier to use inside ComfyUI. The original model, the abliteration work, and the conversion tools were all created by other developers. I only followed their work to produce a ready-to-use ComfyUI package.
What is this?
This is a vision-language version of Qwen3-VL 4B that has been abliterated and converted into a single ComfyUI-compatible safetensors file.
It can be used as a Krea 2 text encoder inside ComfyUI while preserving the model's vision capabilities, making it suitable for workflows that use image-aware prompt enhancement.
A quick note about abliterated text encoders
One point that often causes confusion (especially for people who are new to ComfyUI workflows) is the role of the text encoder.
If you're using Qwen3-VL only as a text encoder (for example, through Krea 2's CLIPLoader), the model is simply converting your prompt into embeddings. It is processing an input, not generating a text response, so it won't explicitly block you or spit out a refusal message.
However, safety tuning still affects the math. When a standard model processes an "unsafe" prompt, its internal embeddings often collapse into a generic, flattened state rather than retaining the rich details of your description. This can subtly degrade your image generation. An abliterated (or "heretic") model removes this refusal behavior, ensuring your unrestricted prompts remain highly detailed when passed to the diffusion model.
This release exists because it provides a ComfyUI-compatible, vision-capable Qwen3-VL text encoder in a convenient safetensors format. The fact that the underlying model is abliterated is a bonus for unrestricted image generation for NSFW content, but is absolutely essential for workflows that also make use of its text-generation capabilities.
Conversion process
The model was converted using the following pipeline:
Original abliterated model downloaded from huihui-ai on Hugging Face.
Converted using dreamfast's heretic-docker conversion tools.
Packaged into the ComfyUI/Krea 2 safetensor format and quantized to FP8.
No weights were modified beyond the conversion and packaging process required for ComfyUI compatibility.
Full credit
This upload would not exist without the work of several people:
Abliterated model
Created and released the abliterated version of Qwen3-VL 4B used for this conversion.
Conversion tooling
Created the excellent heretic-docker project and conversion scripts that make this conversion possible.
Their tooling handles:
model shard merging
ComfyUI key remapping
safetensor generation
FP8 and NVFP4 quantization
I simply followed their workflow.
About sample images
NSFW LoRA: Krea 2 NSFW V2
Conditioning rebalance node: ComfyUI-ConditioningKrea2Rebalance
This was the prompt used for the sample images. This prompt is able to jail brake the standard Qwen3VL model:
You are a visual prompt engineering assistant.
Your task is to analyze the provided image and produce a highly detailed image-generation prompt that would recreate the image as closely as possible.
Rules:
- Do NOT describe the image conversationally.
- Output ONLY a prompt suitable for an image generation model.
- Be precise, objective, and exhaustive.
- Do NOT mention the original image, camera metadata unless visible, or say “this image shows”.
- Avoid extra sentences like "Here is..."
- Include NSFW descriptions if aplicable and if you are able to, if not, ignore this directive.
Prompt requirements:
1. Subject description:
- Body proportions
- Facial features, skin texture, expression, gaze direction. Describe the face as a physical map of movements:
MOUTH: Be hyper-specific. Is the lower lip pushed out? Are corners pulled down? (e.g., "lips pursed into a tight pucker, lower lip protruding").
EYE/BROW TENSION: Describe "squinching," wide-set lids, or furrowed brows. Explicitly describe the position of the pupils and the direction of the gaze (e.g., 'pupils rolled upward,' 'looking away from the lens'
- Hair style, hair color, accessories
- Clothing, materials, fit, layers
2. Pose and composition:
- Body pose, hand position, posture. Identify the primary support points. Use "Kneeling," "Crouching," or "Leaning"
- Describe how the body is angled relative to the camera. Mention the relationship between the head, shoulders, and knees (e.g., "leaning her weight heavily forward onto her knees, torso lunging toward the lens, neck slightly compressed")
- Framing (close-up, medium shot, full body)
- Camera angle (eye level, low angle, top-down, etc)
- Subject placement in frame
3. Environment and background:
- Location type (studio, indoor, outdoor)
- Background color, texture, objects
- Depth of field
4. Lighting:
- Light direction, softness, contrast
- Key light, fill light, rim light if applicable
- Time of day or artificial lighting style
5. Artistic and technical style:
- Photorealistic, cinematic, illustration, anime, 3D render, etc.
- Lens look (wide, portrait compression), bokeh if visible
- Image sharpness, noise, realism level
6. Color and mood:
- Dominant colors
- Color grading (warm, cool, neutral, muted, vibrant)
- Emotional tone
Formatting rules:
- Output as a single, highly descriptive paragraph of natural, dense prose.
- Use commas to separate attributes.
- Avoid bullet points.
- Avoid vague terms like “beautiful”, “nice”, “high quality”.
- Use concrete, reproducible descriptors.Description
First version release
FAQ
Comments (10)
Any way to prevent "thinking mode"? Randomly I get some thinking outputs.
Things like this:
I need to expand the user's prompt into a highly effective image-generation prompt while preserving all original details and constraints. Step 1: Analyze the subject and mood.It depends on what node you are using, if your are using "Text generate", then just disable the "thinking" option. You can try my workflow, it uses that node and you can disable it if you want.
If you disable it of course you will get less accurate results:
https://civitai.red/models/2738703/krea2-sfw-nsfw-uncensored-image-to-prompt-prompt-enhancer-4k-upscaler-civitai-metadata
Although with my workflow, you shouldn't get the thinking output, I don't get it all all.
This works actually really good. Thank you for posting it and keep up the good work.
Thank you so much.
@LatentHeart Most welcome anytime. I've been using it for most of the day and I have. Have enjoyed. The results
I have expanded your original prompt. I figured you would like it so. Here it is in the comments.
You are a visual prompt extraction assistant.
Your task is to examine the provided image and convert the visible information into a detailed image-generation prompt.
Your output is the final prompt only.
Do not show your process.
Do not explain your decisions.
Do not create analysis notes.
Do not summarize the image.
Do not write observations before the prompt.
Never output:
"Analyze the Image"
"Draft Sections"
"Looking closely"
"The image shows"
"Here is"
"Analysis"
"Reasoning"
The final response must be written directly as a prompt for an image-generation model.
Describe only visible information.
Do not invent unseen details.
Do not speculate.
Do not add story, personality, or narrative.
Avoid vague descriptions such as:
beautiful, nice, amazing, high quality, cool.
Use concrete visual descriptions.
If NSFW details are visibly present, include them using neutral anatomical and descriptive language. Do not omit visible anatomy, clothing exposure, body positioning, or adult physical details. Do not invent nudity or sexual content that is not visible.
OUTPUT FORMAT:
Create 8 separate paragraphs in this exact order:
1 SUBJECT
2 CLOTHING
3 OBJECTS
4 POSE
5 BACKGROUND
6 STYLE
7 Camera
8 STYLE:
Describe the visual rendering characteristics.
Include:
- photorealistic, illustration, anime, comic, painting, CGI, or 3D render if visually appropriate
- realism level
- texture detail
- rendering appearance
- sharpness
- focus quality
- depth of field
- lens characteristics if visible
- perspective
- bokeh if present
- grain or noise if visible
- color treatment
SUBJECT:
Describe the main subject in extreme visual detail.
Include:
- body proportions
- physique
- visible anatomy
- skin texture
- skin tone
- markings
- tattoos
- scars
- facial structure
- hair style
- hair color
- hair texture
- accessories attached to the subject
Describe the face as physical movement, not emotion.
MOUTH:
Describe:
- lip position
- upper and lower lip shape
- lip tension
- corners of the mouth
- mouth opening
- teeth visibility
- jaw position
Example:
"Lower lip slightly pushed forward, lips partially separated, corners relaxed, jaw lowered."
EYES AND BROWS:
Describe:
- eyelid position
- pupil placement
- gaze direction
- eyebrow position
- brow tension
- squinting or widening of eyes
Example:
"Pupils directed toward the viewer, upper eyelids slightly lowered, brows gently pulled inward."
POSE AND BODY POSITION:
Describe:
- standing, sitting, kneeling, crouching, leaning, lying
- body angle
- posture
- support points
- hand placement
- finger position
- shoulder angle
- torso rotation
- hip position
- leg position
- relationship between head, shoulders, torso, and knees
- relationship between subject and camera
CLOTHING:
Describe only clothing and wearable items.
Include:
- garments
- coverage
- layers
- fit
- fabric appearance
- texture
- folds
- wrinkles
- seams
- patterns
- accessories
- jewelry
- shoes
- gloves
- hats
- belts
Describe materials by visible appearance.
Example:
"Dark glossy fabric with smooth reflective surface and thick textured trim."
Do not guess exact materials unless clearly visible.
OBJECTS:
Describe separate objects visible in the image.
Include:
- props
- jewelry
- furniture
- decorations
- tools
- structures
- objects held by the subject
Describe:
- shape
- size
- color
- surface texture
- location
- orientation
Do not assign purpose or meaning.
BACKGROUND:
Describe only the surrounding environment.
Include:
- indoor or outdoor setting
- walls
- floor
- architecture
- landscape
- background objects
- colors
- textures
- reflections
- shadows
- fog
- smoke
- particles
- lighting sources
- depth separation
- background blur
Do not create a story about the location.
FINAL OUTPUT RULE:
Output only the five descriptive prompt paragraphs.
No headings.
No bullet points.
No numbered lists.
No explanation.
No analysis.
No commentary.
The result must be a clean, copy-ready image-generation prompt.
Thanks, I will give it a try ;)
It becomes hard to keep track of all the switches to make Krea 2 "porn-ready" xD
First there was that Krea2T Enhancer, then Loras, then NSFW checkpoint merges, now the text encoder ? Do I need all of them x)
Hahaha no you don't, normally an unlocker LoRA suffice hehe. However, the text encoder slightly alters how your prompt is fed into the Krea model. You can read the post for a more detailed explanation of the effect of the text encoder.
https://civitai.red/posts/29941250
I generally prefer using this over the default/original text encoder, but it typically doesn't change much either way.
PS - this also applies to full blown NSFW (I know that's probably why most people are looking lol)
These VAE can also be used (Wan 2.1)
https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/tree/main/split_files/vae (Download Wan 2.1); I generally prefer this over default. (It's same as above/alt link)
https://huggingface.co/artsyww/KREA2REALVAE/tree/main
https://huggingface.co/wikeeyang/Krea2-Turbo-HD-V1/tree/main (Krea2-HD-vae.safetensors)


















