CivArchive
    Model trainer/creator: kblueleaf Uploaded to MoescapeA by Laanver Phi [Please view the full model detail on CivitAI] rev1 Source: https://civitai.com/models/399873?modelVersionId=445973 rev2 Source: https://civitai.com/models/399873?modelVersionId=546178 *Please check the [Usage1] and [Usage2] page to see how to use to checkpoint. Check [Models1] and [Models2] for understanding of the difference between versions. Check [Others] and for other details. Generally check the CivitAI page for complete breakdown. Kohaku' XL εpsilon: The best example of tuning t2i model at home with consumer-level hardware join us: https://discord.gg/tPBsKDyRR5 Introduction: Kohaku--- - XL Epsilon, the fifth major iteration in the Kohaku XL series, features a 5.2 million images dataset, LyCORIS fine-tuning[1], trained on comsumer-level hardware, and is fully open-sourced. Benchmark: *please check the CivitAI page for the graph. CCIP score on 3600 characters. (0~1, higher is better). Clearly, Kohaku XL Epsilon is way better than Kohaku XL Delta

    Description

    Other Training Details: a. Hardware: 4 RTX 3090s b. Num Train Imgs: 5,210,319 c. Total Epoch: 1 - Total Steps: 20354 - Batch Size: 4 - Grad Accumulation Step: 16 - Equivalent Batch Size: 256 d. Optimizer: Lion8bit - Learning Rate: 2e-5 for UNet / 5e-6 for TE - LR Scheduler: Constant (with warmup) - Warmup Steps: 1000 - Weight Decay: 0.1 - β's: 0.9, 0.95 e. Min SNR Gamma: 5 f. Noise Offset: 0.0357 g. Resolution: 1024x1024 h. Min Bucket Res: 256 i. Max Bucket Res: 4096 j. Mixed Precision: FP16 Other Training Details For rev2: a. Hardware: 4 RTX 3090s b. Num Train Imgs: 1,536,902 c. Total Epoch: 5 - Total Steps: 15015 - Batch Size: 4 - Grad Accumulation Step: 32 - Equivalent Batch Size: 512 d. Optimizer: Lion8bit - Learning Rate: 1e-5 for UNet / 2e-6 for TE - LR Scheduler: Cosine (with warmup) - Warmup Steps: 1000 - Weight Decay: 0.1 - β's: 0.9, 0.95 e. Min SNR Gamma: 5 f. Noise Offset: 0.0357 g. Resolution: 1024x1024 h. Min Bucket Res: 256 i. Max Bucket Res: 4096 j. Mixed Precision: FP16 Warning: Versions 0.36.0~0.41.0 of bitsandbytes have significant bugs in the 8bit optimizer that could compromise training, so updating is essential.[8] Training Cost: Utilizing DDP with 4 RTX 3090s, completing 1 epoch across the 5.2 mil image dataset took approximately 12-13 days. Each step for an equivalent batch size of 256 took about 49-50 seconds to complete. Training Cost Rev2: Utilizing DDP with 4 RTX 3090s, completing 5 epoch across the 1.5 mil image dataset took approximately 17-19 days. Each step for an equivalent batch size of 512 took about 105-110 seconds to complete. Why I publish 13600step intermediate ckpt: The training progress have crashed when between 13600step~15300step. And kohya-ss trainer didn't implement resume+step skip before. Athough Kohya and I figured out how to do it correctly and did some sanity check on it. I still cannot fully ensure final result is correct. So I publish the final intermedate ckpt so if anyone want to reproduce training. They have chance to figure the problem of final result.

    FAQ