
Use this with 
this lora has 105 epochs with 850 different images. its better quality for base than other loras and it can do a lot. there is an emphasis on butts for sure but it can do anything I will start posting images since this isn't gaining traction. get the workflow in my article about combination with base and turbo for quality. i
Description
a more refined version with 39,000 steps on 460 images at about 84.7 passes per an image. The dataset featured a highly curated set of 465 images of my ideal feminine beauty in different poses with an emphasis on clothing fit and on and off clothing so that certain fits will land in the model . Essentially curvy and thin women but all with the "bean effect" tight clothing will look tight in the glutes and crotch, faces will be adorable . this is no overfitted and can handle very high and very low strengths due to the large dataset and training steps. I like to use anywhere from 0.5-0.9 but I would recommend trying anything for fun. blend it with other loras and more.
FAQ
Comments (18)
I downloaded and tested your LoRA, and honestly, it still needs a lot of work.
It appears noticeably overtrained right off the bat. The model seems to have learned the colors and tones too aggressively, causing a strange yellowish cast on the output images. This isn't really an issue with the training steps, but rather the quality of your dataset. You should avoid using AI-generated images and instead use high-resolution real photos downscaled to 1536x1536. When training for z-image turbo, 1536px images yield the best results.
Frankly, your recommended weight of 0.8 is practically unusable because it destroys the realism of the base checkpoint. Your LoRA behaves more like a Style LoRA than a proper NSFW LoRA.
I’ve personally trained about 40 LoRAs for z-image, and I've concluded that a minimum of 100 epochs is necessary. Official docs recommend about 200 epochs per image, which implies that a dataset of 450 images would need 45,000(100epochs) to 90,000(200epochs) steps. I usually compromise at around 100 epochs. However, in your case, even training for 90,000 steps would likely result in worse quality simply because the dataset is poor.
As it stands, the LoRA linked below is the best alternative.
https://civitai.com/models/1088938 (11GB lora)
It is actually more of a style lora, It was trained on 1200x800 images. I am aware of the color cast it comes and go's and is the result of a smaller imbedded dataset to add some style to the lora. I have been using it for a while at higher or lower powers and haven't found it to destroy the quality of the base model. But I have found if you use the heretic clip loader it will ruin ruin outputs and make it look overtrained. so I recommend not using the qwen heretic clip loader. I have really tried this at multiple prompts completely on or off and haven't found it too destroy the quality. I'm surprised its giving you that issue! But I am aware that I am using the Q8 gguf and there are a lot of other versions that is way I included a screenshot of my models as loaded in comfyui so people could see how I use it. I do appreciate the feedback though! I'd be curious to know what model version of zimage you run and which clip.
I really prefer the look mine! hahah to each there own. But i'm so happy you made that comparison!
oh holy shit 11gb lora! I must try this
I've love to pick your brain about when I should be using a higher rank?
---
job: "extension"
config:
name: "Quality_1_Training"
process:
- type: "diffusion_trainer"
training_folder: "output"
sqlite_db_path: "./aitk_db.db"
device: "cuda"
trigger_word: ""
performance_log_every: 10
network:
type: "lora"
linear: 128
linear_alpha: 128
conv: 128
conv_alpha: 128
lokr_full_rank: true
lokr_factor: -1
network_kwargs:
ignore_if_contains: []
save:
dtype: "bf16"
save_every: 1000
max_step_saves_to_keep: 16
save_format: "diffusers"
push_to_hub: false
datasets:
- folder_path: "/path/to/images/folder"
mask_path: null
mask_min_value: 0.1
default_caption: ""
caption_ext: "txt"
caption_dropout_rate: 0
cache_latents_to_disk: true
is_reg: false
network_weight: 1
resolution:
- 1024
controls: []
shrink_video_to_frames: true
num_frames: 1
do_i2v: true
flip_x: false
flip_y: false
train:
batch_size: 1
bypass_guidance_embedding: false
steps: 10000
gradient_accumulation: 1
train_unet: true
train_text_encoder: false
gradient_checkpointing: true
noise_scheduler: "flowmatch"
optimizer: "adamw8bit"
timestep_type: "weighted"
content_or_style: "balanced"
optimizer_params:
weight_decay: 0.0001
unload_text_encoder: false
cache_text_embeddings: true
lr: 0.0001
ema_config:
use_ema: false
ema_decay: 0.99
skip_first_sample: true
force_first_sample: false
disable_sampling: true
dtype: "bf16"
diff_output_preservation: false
diff_output_preservation_multiplier: 1
diff_output_preservation_class: "person"
switch_boundary_every: 1
loss_type: "mse"
model:
name_or_path: "Tongyi-MAI/Z-Image-Turbo"
quantize: false
qtype: "qfloat8"
quantize_te: false
qtype_te: "qfloat8"
arch: "zimage:turbo"
low_vram: true
model_kwargs: {}
layer_offloading: false
layer_offloading_text_encoder_percent: 1
layer_offloading_transformer_percent: 1
assistant_lora_path: "ostris/zimage_turbo_training_adapter/zimage_turbo_training_adapter_v2.safetensors"
sample:
sampler: "flowmatch"
sample_every: 250000
width: 1024
height: 1024
samples:
- prompt: "man with red hair, playing chess at the park, bomb going off in the background"
- prompt: "a man holding a coffee cup, in a beanie, sitting at a cafe"
- prompt: "a horse is a DJ at a night club, fish eye lens, smoke machine, lazer lights, holding a martini"
- prompt: "a man showing off his cool new t shirt at the beach, a shark is jumping out of the water in the background"
- prompt: "a bear building a log cabin in the snow covered mountains"
- prompt: "man playing the guitar, on stage, singing a song, laser lights, punk rocker"
- prompt: "hipster man with a beard, building a chair, in a wood shop"
- prompt: "photo of a man, white background, medium shot, modeling clothing, studio lighting, white backdrop"
- prompt: "a man holding a sign that says, 'this is a sign'"
- prompt: "a bulldog, in a post apocalyptic world, with a shotgun, in a leather jacket, in a desert, with a motorcycle"
neg: ""
seed: 42
walk_seed: true
guidance_scale: 1
sample_steps: 8
num_frames: 1
fps: 1
meta:
name: "[name]"
version: "1.0"
It seems it simply comes down to a difference in our aesthetic preferences. I've tested nearly every z-image checkpoint released on Civitai and other forums, but I have yet to find a model that can replace the official one. Most of them are still garbage.
I currently use z-image turbo bf16 with the standard VAE and text encoder. Occasionally, I use the abliterated text encoder shared by huihui when my prompts get blocked: https://huggingface.co/huihui-ai/Qwen3-4B-abliterated
I recommend using Rank 128. I’m sharing my ostris training settings with this message. I hope this helps. but it's for 24gb gpu setting.
@RABBAI I agree the alternative checkpoints are all garbage. Its a waste of time to make them until the true base model is out. thank you for the information. I tried to figure out what was best for getting the details I wanted. I saw a jump in quality when I went from rank 32 to 64 but was worried about overdoing it and wasting hrs of runpod time so that's why I jump up to 64 rather than 128. There is a lot of people on youtube that do tutorials on lora training that clearly know no more than me which is frustrating. I will likely take another stab at this lora in the future, I'll run rank 128 and perhaps I'll prune the style elements out of the dataset. I'll likely also replace all the photos with best quality versions but that will be a task ugh.
@iamgeekusa Yes, I look forward to your new model. However, don't spend too much time on it. I used to train a ton of Flux models, but I found that they all became useless over time. Who knows, maybe some crazy company will eventually release a fully NSFW-capable model.
By the way, the Ostris preset I shared is set to 1024px, but you need to change that to 1536px. I must have sent the one I was using for some recent tests with a 1024 dataset. Cheers.
@RABBAI he did after all call it the bean effect
@RABBAI I think if I make another I might lean into it being even more of a style lora and a NSFW lora. Mystic xxx is clearly better for just basic stuff. Its kinda why I did mine anyway. Although if you mix mine with mystic xxx you can get some fun stuff. mystic .60 mine 40
No offense, but in your comments here you just come across as a bit too pretentious to me.
Oh well, whatever floats your fancy I suppose.
@Latent_Dreamscape really? I was just happy to get any feedback not sure how I'm being pretentious. oh well.
you should try my new one with the workflow in my newest article, I think you'll be impressed
"You won't get cuter looking woman out of any other lora" I dont know there are some contenders but i keep those a secret :3 lol I love the name you chose very fitting for your project and i see similar things i like in what you have done so i'll give it a try thanks :D
yea that was a stretch ahah but tastes vary, I'm not a fan of the current AI thick eyebrows lady people keep training that seem to be pony based
yeah I wondered how man people would get the crumb reference
we can see your images bro it's not a secret ":3"







