CivArchive
    CyberRealistic Z-Image Turbo - v5.0
    NSFW
    Preview 130909706
    Preview 130909693
    Preview 131001861
    Preview 130909687
    Preview 130909641
    Preview 130909642
    Preview 130909645
    Preview 130909648
    Preview 130909653
    Preview 130909688
    Preview 130909694
    Preview 130909716
    Preview 130909724

    You can get this model through Civitai Early Access, or grab it along with many others by joining The Tinkerer on Whop. Membership gets you early releases, private tools and members-only pages. There are free pages too, no membership needed.

    👉 Join on Whop
    💬 Join the community for support, free tools and early news on Discord


    CyberRealistic Z-Image Turbo is a realism-focused finetune of Z-Image Turbo by Tongyi-MAI.

    The idea behind it is deliberately simple: keep what makes Z-Image Turbo good - speed, strong prompt understanding, good composition and extremely efficient few-step generation - while moving the default visual language further toward believable photography.

    CyberRealistic doesn't try to turn Z-Image Turbo into a completely different model. The original already has a very capable photographic foundation. The finetune mainly changes what the model considers a "normal" photograph: more natural skin, less synthetic rendering, more believable faces, stronger material texture, more grounded lighting and better anatomical consistency.

    Z-Image Turbo already knows how to make a good image. CyberRealistic mainly changes where it starts.

    What's different from base Z-Image Turbo

    • Stronger photographic look out of the box.

    • More natural skin texture with less waxy or overly polished rendering.

    • Improved faces, eyes, hair and small facial details.

    • Better anatomical consistency, especially hands, feet and complex poses.

    • Fewer duplicated limbs and extra hands in more difficult compositions.

    • More believable fabric, hair, skin, metal, glass and other material textures.

    • Stronger response to available light, practical lighting and real-world camera language.

    • Less dependence on stacks of words like masterpiece, 8k, ultra detailed and photorealistic.

    • Keeps the speed and general prompt behavior that make Z-Image Turbo useful.

    The focus is photography, but that doesn't mean the model is locked to photography. Illustration, cinematic stylization, fantasy, advertising, vintage photography and other looks are still available when you describe them.


    Prompting

    If you're coming from SDXL, Pony or Illustrious, the biggest change is simple:

    Describe the image instead of building a tag stack.

    Z-Image Turbo uses a Qwen3-based text encoder and responds very well to normal descriptive language. Short comma-separated clauses are completely fine, but every part of the prompt should ideally tell the model something visual.

    Instead of:

    woman, realistic, masterpiece, best quality, detailed skin, cinematic, 8k
    

    try:

    A woman standing beside an open apartment window on a warm summer evening, photographed with soft natural light falling across her face, loose dark hair, natural skin texture and an out-of-focus city street behind her.
    

    The second prompt gives the model an actual scene to construct.

    Put the subject first

    Start with what the image is about.

    A middle-aged mechanic leaning over the open engine bay of an old red pickup truck...
    

    works better than hiding the subject halfway through a long list of style instructions.

    You don't need to obsess over exact prompt order, but the main subject and composition should be clear early.

    Be specific

    Specific visual language usually does more than generic quality words.

    Instead of:

    beautiful lighting
    

    try:

    soft late-afternoon sunlight entering through a dusty workshop window
    

    Instead of:

    detailed clothing
    

    try:

    a faded blue denim jacket with worn seams and slightly frayed cuffs
    

    Instead of:

    cinematic portrait
    

    try:

    photographed from chest height with a 50mm lens, shallow depth of field and soft window light from camera left
    

    Describe the light

    Lighting is one of the easiest ways to change the realism and mood of the image.

    Useful examples:

    soft overcast daylight
    
    direct midday sunlight creating hard shadows
    
    a single warm tungsten lamp above the table
    
    cold fluorescent supermarket lighting
    
    late-afternoon sunlight entering through venetian blinds
    
    direct on-camera flash in a dark room
    

    You can still use words like cinematic, but describing where the light actually comes from gives the model much more information.

    Quality tags are not magic switches

    Words such as:

    masterpiece
    best quality
    8k
    ultra detailed
    absurdres
    score_9
    

    can still influence the wording of the prompt, but Z-Image Turbo doesn't treat them like the traditional SDXL/Pony quality system.

    Use that prompt space to describe what you actually want to see.

    Camera language works well

    For photographic images, camera terminology can be useful when it describes a visible effect:

    35mm documentary photograph
    
    85mm portrait lens with shallow depth of field
    
    handheld photograph with slight motion blur
    
    direct flash snapshot
    
    wide-angle environmental portrait
    
    medium-format color photograph
    

    Don't feel forced to specify a camera and lens in every prompt. Sometimes simply saying casual phone photo gives you exactly the look you need.

    Prompt length

    There is no perfect prompt length, but these are useful practical ranges:

    • 10–30 words: exploration and seed hunting.

    • 30–80 words: good balance between control and freedom.

    • 80–150 words: complex scenes, precise lighting or detailed compositions.

    Long prompts aren't automatically better. Contradictory prompts are the bigger problem.

    If you ask for soft natural window light, hard direct flash, deep cinematic shadows and flat commercial studio lighting at the same time, the model still has to decide which instruction wins.

    Text inside images

    Z-Image Turbo is unusually capable at rendering text compared with older diffusion models.

    If exact text matters, put it in quotation marks:

    A small neon sign above the diner entrance reading "OPEN ALL NIGHT"
    

    Keep important text reasonably short. It's good, but it still isn't a replacement for a typography application.


    Z-Image Turbo is a distilled few-step model.

    Don't treat it like an SDXL checkpoint that needs 30–50 steps.

    A good starting point is:

    • Steps: 8–9

    • CFG / Guidance: effectively OFF

    • Resolution: start around 1 megapixel and increase if your hardware allows it

    • Negative prompt: normally unnecessary

    In the original Diffusers implementation, guidance is 0.0.

    In standard ComfyUI workflows, the equivalent no-CFG setup is generally CFG 1.0.

    More steps are not automatically better with Turbo. If something isn't working, changing the prompt, seed, sampler or composition usually makes more sense than simply increasing the step count.

    ComfyUI

    For ComfyUI, I recommend starting with the current Z-Image Turbo workflow/template rather than applying old SDXL settings.

    Z-Image Turbo has its own sampling behavior and is designed around very low step counts.

    Negative conditioning is normally zeroed out in the standard Turbo workflow because the model runs without traditional classifier-free guidance.


    Example prompts

    Natural-light portrait

    A woman in her early thirties sitting beside an open café window, loose brown hair falling across one side of her face, wearing a simple cream-colored sweater. Photographed from slightly below eye level with a 50mm lens, soft overcast daylight entering from the window, natural skin texture, muted colors and a busy street softly blurred in the background.
    

    Documentary photography

    An elderly fishmonger arranging silver mackerel on crushed ice at an indoor market early in the morning. Cold daylight enters through the open market doors and mixes with the warm bulbs above the counter. Wet concrete floor, weathered hands, faded rubber apron, handheld 35mm documentary photograph with subtle grain and natural color.
    

    Low-light snapshot

    A young woman standing alone beside a vending machine outside a convenience store at two in the morning, photographed with direct on-camera flash. Dark parking lot behind her, slightly messy hair, casual oversized jacket, realistic skin texture, hard flash shadows, muted colors and the imperfect look of a spontaneous late-night photograph.
    

    A few last things

    Short prompts are completely valid.

    One of the advantages of Turbo is that you can generate several directions quickly, choose the seed or composition you like, and then add more camera, lighting and material detail.

    That often works better than trying to write the perfect 150-word prompt before generating anything.

    Also keep in mind that Z-Image Turbo is distilled for speed. Part of that tradeoff is lower variation than a large non-distilled foundation model. If you keep seeing the same interpretation, change the wording more substantially rather than adding another five quality tags.

    CyberRealistic Z-Image Turbo is released for people who enjoy generating, experimenting, benchmarking and finding the edges of a model.

    Feedback is especially useful for difficult poses, multiple people, hands and feet, unusual lighting, text rendering and prompts where the model behaves differently from the original Z-Image Turbo.

    If you find something interesting — good or bad — let me know.

    Credits

    CyberRealistic Z-Image Turbo is based on Z-Image Turbo by Tongyi-MAI.

    Z-Image Turbo is released under the Apache 2.0 License. Please follow the applicable upstream license when using or redistributing derived models.

    Description

    Join The Tinkerer on Whop. Membership gets you early releases, private tools and a bunch of extra stuff.
    👉 Join on Whop

    V5 refines what V4.0 started. Better details, sharper colours, more consistent output, and stronger NSFW content across the board.

    This release adds 110+ new sample images, covering a wider range of styles and scenarios. All images were generated with the Cyber Z-Image Turbo Workflow v4.1, available in the Optional Files section.

    FAQ

    Comments (47)

    vicautMay 16, 2026· 5 reactions
    CivitAI
    Cyberdelia
    Author
    May 16, 2026

    @vicaut looks promising. Will test it.

    Beezer79May 16, 2026

    thanks. will test it, too.

    what's this for? what does it do? how does it compare to joshepe?

    how does one make the two files work? i has 1 of 2 and 2 of 2.

    Cyberdelia
    Author
    May 16, 2026· 1 reaction
    Beezer79May 16, 2026

    @Melodic_Possible_582589 i used the gguf Q8 version https://huggingface.co/BennyDaBall/Qwen3-4b-Z-Image-Engineer-V4/tree/main

    BzzzDarklordMay 16, 2026

    Thank you for your contribution to our community. hands on

    vicautMay 16, 2026

    @Melodic_Possible_582589 Join them with https://github.com/soursilver/safetensors-merger

    vicautMay 16, 2026

    @Cyberdelia Try also Euler ancestral combined with FlowMatch Euler Discrete Scheduler. This is my all time favoutite combination for z image turbo.

    Cyberdelia
    Author
    May 17, 2026

    @vicaut Ah, I thought it was a text encoder. I use a custom text encoder for this myself. And for creating Z-Image prompts, I also use a different system. This is fine in itself - I was just a bit confused about what it actually was. It was also very late last night :)

    vicautMay 17, 2026

    @Cyberdelia Well, it is? Z-Engineer V4 is a fully fine-tuned version of the text encoder from Tongyi-MAI/Z-Image-Turbo. It's been specifically trained to understand the nuances of AI Image Generation workflows.

    It excels at:

    Expanding Concepts: Turn "sad robot in rain" into a cinematic fever dream with chromatic aberration, shallow depth of field, and a melancholic color grade that would make Blade Runner jealous.

    Technical Precision: It knows the difference between an 85mm portrait lens and a 24mm wide—and will use them appropriately. Lighting? Rembrandt, split, volumetric fog? It's got opinions.

    Stylistic Consistency: It writes with a creative voice, not that robotic "hyperrealistic, 8k, trending on artstation" energy.

    vicautMay 17, 2026

    @Melodic_Possible_582589 It is a better, fune-tuned, text encoder.

    Cyberdelia
    Author
    May 17, 2026· 1 reaction

    @vicaut Aha, so it is a text encoder after all :) Then I’ll compare it with my own text encoder.

    @vicaut thanks for the info. I will have to try it again. When I compared it to the josiefied version the z-engineer couldn't do penetration in some prompts. I also used both the both the 8 and 16 bit version of z-engineer in LM studio and it couldn't really produce good zimage style prompts. I was able to reach the realism generations because I used Cyberdelia's prompt on chatgpt, but that doesn't allow nsfw, so I describe the character on cyberdelia's program and describe the nsfw stuff using LM studio. I have not used lm studio for awhile now.

    Cyberdelia
    Author
    May 18, 2026· 3 reactions

    @vicaut @Beezer79 @Melodic_Possible_582589 I've updated my workflow and created a new ComfyUI node based on BennyDaBall930's original ComfyUI-Z-Engineer.
    https://civitai.red/models/2532359/cyberrealistic-z-image-turbo-comfyui-workflow?modelVersionId=2957140

    Tofu080May 17, 2026
    CivitAI

    any GGUF version 🥲?

    ferretduckMay 17, 2026· 3 reactions

    @Tofu080 I made an open source easy-to-use program with low RAM usage so you can make your own GGUF conversions: https://github.com/qskousen/ggufy

    Cyberdelia
    Author
    May 19, 2026· 1 reaction

    @ferretduck yes, works perfect!

    dillion1920Jun 1, 2026· 1 reaction

    Pretend i am an imbecile -not with too much enthusiasm though!- but why do you want gguf?

    Cyberdelia
    Author
    Jun 1, 2026

    @dillion1920 good question and I have no idea!

    Tofu080Jun 2, 2026· 1 reaction

    @dillion1920 because i dont have enought vram for it, in order for it to work i need to offload to my system Ram. i found this tool and it did great job https://github.com/SlaveOfGod1/ggufy

    Cyberdelia
    Author
    Jun 2, 2026· 1 reaction

    @Tofu080 Ah got it.

    ferretduckJun 2, 2026

    @Tofu080 wow, interesting! i didn't know anyone had forked ggufy. it seems somewhat limited so far, but interesting that they tried to rewrite it in python

    dsanatlarJun 4, 2026

    @ferretduck What are the minimum specs for doing this?

    ferretduckJun 17, 2026

    @dsanatlar you just need a CPU and a couple gigabytes of RAM, depending on the model you want to convert.

    svyatkin_pMay 19, 2026· 1 reaction
    CivitAI

    Any difference in input betwen fb8 and bf16?

    ss9999Jun 2, 2026· 1 reaction

    bf16 has higher quality because it's less compressed.

    mphobbitMay 20, 2026· 4 reactions
    CivitAI

    Check please the checkpoint on Civit generator, it gives only digital noise for some reason.

    Cyberdelia
    Author
    May 20, 2026· 1 reaction

    I didn't know that it was possible to use it in the generator. This is a new thing, and I don't know why it doesn't work. Let me check with Civitai about this.

    danedina432May 25, 2026· 1 reaction

    same

    amazingbeautyJun 1, 2026· 2 reactions
    CivitAI

    FP16 needed.

    from SDXL, cyberrealistic is a details machine, but anything else? it depends..

    SovietmadeJun 1, 2026· 7 reactions
    CivitAI

    Kissing your hands, maestro. Works brilliant (currently testing on ref images, gonna try with my custom lora once trained) 🤌🤌🤌

    6vidit9Jun 1, 2026· 6 reactions
    CivitAI

    V5 is absolute CINEMA

    olternautJun 2, 2026· 4 reactions
    CivitAI

    This version 5 is the best z-image-turbo model so far.

    ArbellazJun 3, 2026· 1 reaction
    CivitAI

    This is a massive improvement over the default model (which already looked great). Really nice composition and characters.

    YuriyensJun 5, 2026· 3 reactions
    CivitAI

    It does not work right on civitai, FYI

    jomefarit14230Jun 5, 2026· 1 reaction
    CivitAI

    Something is ruining many pictures across all ZIT models, until a resourceful creator can fix it with a LoRA. Whenever a woman stands full frontal to the camera, her pussy crack reaches too high. A few seconds watching real photos (in sites like www.metarthunter.com or similar) will very quickly make you see the problem as I see it.
    Another weak point of ZIT is nipples that disappear behind even a light blouse or t-shirt, instead of being partially noticed.

    TheodorSidJun 6, 2026· 2 reactions
    CivitAI

    What about non-turbo version ? I can't live without negative ..

    Cyberdelia
    Author
    Jun 7, 2026· 2 reactions

    Haha, I know you like to stay on the negative side of life. 😉

    There is a Base version, but honestly I haven’t had great results with it so far. That’s why I’m mainly focusing on Turbo right now. For me it gives better results and is much more fun to work with.

    t2kardyyJun 6, 2026· 2 reactions
    CivitAI

    Is there already a good way to train a LoRA specifically for this? When using LoRAs that I trained for z-image base, the results are underwhelming. I have also tested inference with the catalyst model and I have not found it to be performing much better.
    Maybe someone made a training-adapter so we can train LoRAs directly on this checkpoint?

    Other than difficulties with LoRAs the V5 model is absolutely insane. Thank you for your work. I greatly appreciate what you do.

    StaGerJun 6, 2026· 3 reactions
    CivitAI

    It works, but not on Civitai for some reason. A pity, because the model is very nice.

    soyv4Jun 6, 2026· 2 reactions
    CivitAI

    The model is really good, but there are exaggerations of color in it, I mean that the color is greatly overestimated for realism, in general, I liked the model, and the textures and understands the norm of promt. I would like to see a little less bright contrasting colors in the next update so that the model can draw a little more realistically. I use dpmpp_2s_ancestral+ (beta57/bong_tangent) or res_2s+ (beta57/bong_tangent) for realism.

    Cyberdelia
    Author
    Jun 7, 2026· 1 reaction
    soyv4Jun 7, 2026· 1 reaction

    @Cyberdelia Unfortunately, I didn't like the Catalyst model.

    Cyberdelia
    Author
    Jun 12, 2026· 3 reactions

    @soyv4 This weekend I will release V6.0 -> The color volume has been slightly toned down with some other improvements

    9scoreJun 7, 2026· 14 reactions
    CivitAI

    Best nsfw/female model by far 10/10

    Checkpoint
    ZImageTurbo

    Details

    Downloads
    6,614
    Platform
    CivitAI
    Platform Status
    Available
    Created
    5/16/2026
    Updated
    10/5/2026
    Deleted
    -

    Files

    ae.safetensors

    Mirrors

    HuggingFace (792 mirrors)
    TensorFiles (1 mirrors)
    ShakkerAI (1 mirrors)

    cyberrealisticZImage_v50_txt.safetensors

    Mirrors

    HuggingFace (269 mirrors)

    cyberrealisticZImage_v50.zip

    Mirrors

    CivitAI (1 mirrors)