CivArchive
    Zanime - v1.0
    NSFW
    Preview 1
    Preview 2
    # ๐Ÿ”ฌ Z-Image Base Anime Finetuning โ€“ Full Technical Test Report
    Epoch 100 Evaluation
    
    This guide documents a complete testing and evaluation process of a 
    Z-Image Base anime finetuning checkpoint, including training details, 
    inference settings, prompt engineering findings, and sampler recommendations.
    
    All findings are based on real testing with Epoch 100 checkpoint.
    
    ------------------------------------------------------------
    ๐Ÿง  TRAINING DETAILS
    ------------------------------------------------------------
    
    Base Model:        Z-Image Base (Tongyi-MAI, released Jan 27, 2026)
    Architecture:      S3-DiT (Single-Stream Diffusion Transformer)
    Text Encoder:      Qwen-based (bilingual EN/CN)
    Training Type:     Checkpoint Finetuning (not LoRA)
    Epochs:            100
    Steps:             68,830
    Dataset Size:      1,375 Anime Images
    Tagging System:    WD Tagger (Booru-style tags)
    
    Avg Tags/Image:    ~47 tags
    Unique Tags:       4,834
    Total Tag Count:   64,276
    
    
    ------------------------------------------------------------
    ๐Ÿ“ DATASET RESOLUTION DISTRIBUTION
    ------------------------------------------------------------
    
    Resolution   | Count | Ratio   | Quality
    -------------|-------|---------|---------
    1344x1728    | 132   | 3:4     | โœ… Good
    768x1086     | 95    | ~9:13   | โš ๏ธ Odd-Size
    832x1216     | 47    | ~2:3    | โœ… SD-Standard
    1152x1536    | 45    | 3:4     | โœ… Good
    768x1084     | 36    | Odd     | โš ๏ธ Odd-Size
    768x1024     | 28    | 3:4     | โœ… Perfect
    896x1152     | 22    | 7:9     | โœ… Good
    1365x768     | 20    | ~16:9   | โ†”๏ธ Landscape
    1248x1824    | 19    | ~2:3    | โœ… Good
    768x768      | 19    | 1:1     | โœ… Standard
    
    
    ------------------------------------------------------------
    ๐Ÿท TOP 50 TRAINING TAGS
    ------------------------------------------------------------
    
     1. 1162x  1girl
     2. 1034x  looking_at_viewer
     3. 1033x  solo
     4. 1001x  
     5.  980x  long_hair
     6.  839x  blush
     7.  689x  smile
     8.  595x  large_
     9.  529x  long_sleeves
    10.  520x  closed_mouth
    11.  466x  open_mouth
    12.  456x  bare_shoulders
    13.  426x  hair_between_eyes
    14.  422x  shirt
    15.  420x  thighs
    16.  417x  blue_eyes
    17.  380x  
    18.  376x  medium_
    19.  370x  short_hair
    20.  354x  hair_ornament
    21.  344x  black_hair
    22.  340x  collarbone
    23.  328x  dress
    24.  327x  simple_background
    25.  317x  jewelry
    26.  308x  holding
    27.  299x  indoors
    28.  298x  navel
    29.  297x  sitting
    30.  285x  outdoors
    31.  284x  standing
    32.  282x  gloves
    33.  275x  skirt
    34.  270x  very_long_hair
    35.  269x  jacket
    36.  269x  white_background
    37.  268x  animal_ears
    38.  259x  brown_hair
    39.  253x  blonde_hair
    40.  236x  thighhighs
    41.  232x  white_shirt
    42.  225x  red_eyes
    43.  220x  parted_lips
    44.  219x  multicolored_hair
    45.  216x  cowboy_shot
    46.  214x  bow
    47.  214x  sky
    48.  214x  sweat
    49.  207x  ribbon
    50.  207x  purple_eyes
    
    ------------------------------------------------------------
    ๐Ÿท TOP 50 TRAINING TAGS 
    ------------------------------------------------------------
    
     1.    139x  
     2.    120x  
     3.    117x  
     4.    111x  
     5.    100x  from_behind          
     6.     98x  lying                
     7.     94x  
     8.     81x  
     9.     80x  covered_      
     10.    73x  completely_
     11.    72x                    
     12.    70x  
     13.    67x  _visible_through_thighs
     14.    65x  spread_legs          
     15.    64x  
     16.    61x  _juice
     17.    60x  
     18.    59x  saliva               
     19.    57x  
     20.    53x  
     21.    52x  
     22.    50x  pov                  
     23.    45x  _from_behind
     24.    44x  _from_behind
     25.    41x  huge_         
     26.    40x  
     27.    39x  clothed_
     28.    38x  
     29.    36x  
     30.    35x  bent_over            
     31.    34x  wet_clothes          
     32.    33x  oral
     33.    32x  straddling           
     34.    31x  no_
     35.    31x  _apart
     36.    30x  _grab
     37.    29x  _in_
     38.    29x  
     39.    27x  
     40.    27x  rolling_eyes         
     41.    26x  yuri
     42.    25x  
     43.    25x  _out
     44.    24x  _only
     45.    23x  
     46.    22x  standing_
     47.    22x  cleft_of_venus
     48.    22x  
     49.    22x  
     50.    21x  _overflow
    ------------------------------------------------------------
    โš™ INFERENCE SETTINGS โ€“ WHAT WORKS
    ------------------------------------------------------------
    
    Recommended Setup:
    
    CFG:               4 โ€“ 6  (sweet spot confirmed)
    Steps:             30 โ€“ 40
    Resolution:        768x1024 (primary)
                       832x1216 (more detail)
    ModelSamplingFlow: Shift 3.0  โ† important
    CFG Normalization: NOT tested 
    
    ------------------------------------------------------------
    ๐ŸŽ› SAMPLER & SCHEDULER RESULTS
    ------------------------------------------------------------
    
    CONFIRMED WORKING (anime-style output):
    
    โœ” Euler Ancestral  + Simple
    โœ” Euler Ancestral  + Normal
    โœ” DPM++ 2M         + Simple
    โœ” DPM++ 2M         + Normal
    โœ” DPM++ 2M SDE     + Simple
    โœ” DPM++ 3M SDE     + Simple
    โœ” Res Multistep    + Simple
    โœ” Res Multistep    + Normal
    
    COMPLETELY BROKEN (unrecognizable output):
    
    โœ˜ All Karras variants
    โœ˜ All Exponential variants
    
    Notes:
    
    โ†’ DPM++ 2M SDE and DPM++ 3M SDE tend to produce more realistic-looking backgrounds
    โ†’ All 8 working samplers produce top quality results
    โ†’ Personal preference decides final choice
    
    ------------------------------------------------------------
    ๐Ÿงช PROMPT ENGINEERING FINDINGS
    ------------------------------------------------------------
    
    WD Tags (Booru-style):
    + Fast to write
    + Good character details
    + Good clothing recognition
    - Slightly flatter clothing textures
    - Less atmospheric backgrounds
    - Less "alive" feeling overall
    
    Fulltext English:
    + Richer clothing details and textures
    + Better atmospheric backgrounds
    + More dynamic and "alive" feeling
    + Utilizes Qwen encoder strength fully
    + Better scene composition
    - Slightly longer to write
    
    ------------------------------------------------------------
    ๐Ÿ† WINNING PROMPT STRUCTURE โ€“ LAYERED FULLTEXT
    ------------------------------------------------------------
    
    1. Opening line โ€“ Subject + Style
    2. Character details โ€“ Clothing + Features
    3. Action + Pose
    4. Foreground + immediate environment
    5. Background description
    6. Composition + Lighting + Meta
    
    ------------------------------------------------------------
     NEGATIVE PROMPT FINDINGS
    ------------------------------------------------------------
    
    Rule:
    POSITIVE โ†’ Fulltext
    NEGATIVE โ†’ Short keyword tags
    
    ------------------------------------------------------------
    ๐Ÿ”ค TEXT GENERATION CAPABILITY
    ------------------------------------------------------------
    
    Status after finetuning: INTACT โœ…
    
    Tested:
    โœ” Comic book covers with title text
    โœ” "BLADE ZERO" title text
    โœ” "ANIME MONTHLY" magazine cover
    โœ” Issue numbers and dates
    
    Notes:
    โ†’ Large text works very well
    โ†’ Small text slightly blurry (base limitation)
    โ†’ Occasional spelling errors (base model behavior)
    
    ------------------------------------------------------------
    โš  KNOWN LIMITATIONS
    ------------------------------------------------------------
    
    :
    โ†’ Extra fingers / malformed hands still occur
    โ†’ Floating limbs appear occasionally
    โ†’ Manageable with negative prompts
    โ†’ Known Z-Image Base issue, not training fault
    
    Style Consistency:
    โ†’ Base model produces anime style ~50% of the time
    โ†’ Finetuned model produces anime style consistently โœ…
    
    Details:
    โ†’ Best detail at CFG 5โ€“6, Steps 35โ€“40
    โ†’ ModelSamplingFlow Shift 3.0 is essential
    โ†’ Without Shift results are noticeably worse
    
    
    ------------------------------------------------------------
    ๐Ÿš€ QUICK START SETTINGS
    ------------------------------------------------------------
    
    Node:         ModelSamplingFlow โ†’ Shift 3.0
    Sampler:      DPM++ 2M SDE  or  Euler Ancestral
    Scheduler:    Simple
    CFG:          5
    Steps:        35
    Resolution:   768x1024
    Prompt style: Layered Fulltext
    Negative:     Short keyword tags

    Description

    Release Version with >100 hours computing time on a RTX5090

    Checkpoint
    Z Image Base

    Details

    Downloads
    74
    Platform
    SeaArt
    Platform Status
    Available
    Created
    3/12/2026
    Updated
    3/3/2026
    Deleted
    -

    Files

    Available On (1 platform)

    Same model published on other platforms. May have additional downloads or version variants.