CivArchive
    Krea-2 raw Lora training guide with RTX 3060 12 GB VRAM + 32 GB RAM using Ostris AI Toolkit - v1.0
    NSFW
    Preview 137226902

    This is a short guide how you can train the brand new Krea-2 raw Loras in high quality with low budget hardware by using Ostris AI Toolkit.

    The created Krea-2 raw Loras you can use with Krea-2 raw models as well with the much faster Krea-2 Turbo model.

    By using the Turbo model with your trained raw Lora you can generate images very fast and in good quality with 12 GB VRAM too. This works perfectly with comfyui for example.

    I struggled for several days with Musubi Trainer and other tools to get it running on my low budget system, but it all did not work well.

    Finally I found the the main parts for a very easy to use solution with Ostris AI Toolkit here: reddit (all credits go to the author).

    Because this works very well after some testing and it produces high quality results I would like to publish it here with all setting parameters.

    Ok - lets go:

    What do you need:

    • RTX 3060 GPU with 12 GB VRAM,

    • 32 GB RAM,

    • probably a swap-file on a fast SSD: for myself I use a very large swap-file size of 64 - 128 GB (for video generation), but I`m pretty shure it will work with a much smaller size, like 32 GB for example.

    • Ostris AI Toolkit (github-page). On Windows just use Travis1 Easy-Install.

    • my settings (use the atteched config.yaml file or the screenshots).

    • around 20 - 30 high quality images for a character Lora for example.

    How long will the training run:

    A complete character training for a high quality Lora with 750 training steps will take around 8 - 9 hours. So this works well as an overnight job.

    What are the limitations with 12 GB VRAM:

    • the training resolution is limited to 512+768. 1024 work too, but training will need several days. So 512+768 works well and delivers very good results with Krea-2 Turbo image generation.

    • rank 24 insted of 32 (I don't think it makes much of a difference).

    • Learning Rate = 0,0002 insted of 0,0001 to reduce the steps (training time to 50%) - see below.

    Preperation:

    Training Images (aspect ratio, resizing):

    Use congruent, sharp high quaility images (at least 1024 x 1024) of any aspect ratio. Do not downscale your images. AI Toolkit creates automatically buckets.

    Captions:

    Character training works well without any captions. But if you like, you can use captions too.

    How to start:

    Open AI Tollkit and create a new Dataset with your images (and optional with your captions). Create a new Job and use my settings from config.yaml or the screenshot.

    For a Dataset without captions don`t change anything exept:

    • set your Trigger Word,

    • select your Target Dataset,

    For a Dataset with captions don`t change anything exept:

    • delete the Trigger Word and leave it blank,

    • select Cache Text Embeddings (Unload TE gets unselected),

    • select your Target Dataset.

    Some hints:

    • AI Toolkit automatically downloads the correct Krea-2 raw model files after starting the job (make sure you have sufficient free disk space).

    • AI Toolkit do Cache Latents and Unload Text Encoder automatically to save VRAM for training.

    • We use 750 steps and Gradient Accumulation = 2 - this means we train with 1500 "virtual" steps.

    • With Leraning Rate = 0,0002 we get well trained character Loras mostly between 400 and 750 steps. Because we save every 100 steps you finally can choose the best output.

    • Do not activate Sampling. This leads to OOM errors.

    • The first training steps take very long (around > 300 seconds). Don`t worry. It speeds up with every step and the complete training will take around 8 to 9 hours.

    Description

    FAQ

    Comments (22)

    EvilPeterPanJul 20, 2026· 1 reaction
    CivitAI

    is there any settings i can use to get it to work? (RTX 4070 laptop gpu 8gb, 16gm ram)

    arkinson
    Author
    Jul 20, 2026

    @EvilPeterPan Uhh - 8 GB VRAM + 16 GB RAM - I believe this is actually too much underpowered. You might try lowest settings as possible: resolution <= 512 (256), rank <= 16, learning rate = 0,0003, steps <= 500. Finally it could work, but I would bet, this will not be much fun 😢

    Another question is whether you can actually generate Krea-2 Turbo images properly with your hardware?

    EvilPeterPanJul 20, 2026

    @arkinson don't know haven't tried yet

    arkinson
    Author
    Jul 20, 2026

    @EvilPeterPan Yeah, that`s the reason I asked for.

    EvilPeterPanJul 20, 2026· 1 reaction

    @arkinson i've tried it and i can make images using krea 2

    album3826Jul 20, 2026· 1 reaction
    CivitAI

    Would you change any settings for RTX 5080 16gb VRAM and 32 gb RAM? Also is there any way to avoid downloading the raw model if I already have them? (Assuming I can use my int8 convrot model for training)

    arkinson
    Author
    Jul 20, 2026

    @album3826 Just start with my settings, it should run much faster on your machine. Then try the common tweaks: increase resolution, rank and steps (reduce learning rate). But I`m afraid the main bottleneck will be your 32 GB ram.

    Maybe it is possible to use pre-installed/"custom" models. But for Lora training I definitely would use the full raw model.

    album3826Jul 20, 2026

    @arkinson I see, thanks! The model i was talking about is the raw int8 convrot from here: https://huggingface.co/Comfy-Org/Krea-2/tree/main/diffusion_models, not some weird custom checkpoint. But you're saying it should be the full raw bf16 model?

    arkinson
    Author
    Jul 20, 2026

    @album3826 Sorry, I can`t help you with manual model installation. You have to check AI Toolkit by yourself for integration, format, model type, model path conventions, model management, etc.

    I rather would go the most eseast way: just select Krea-2 raw according to my settings and let AI Toolkit do the complete installation job in the background.

    album3826Jul 20, 2026· 1 reaction

    @arkinson np I did the same, currently training eager to see results!

    arkinson
    Author
    Jul 21, 2026

    @album3826 Good Luck!

    And just a hint: I believe, even with Krea-2, the best way is to train all Loras with the standard model. You will get more flexibillity. From my short tests I allready can see, that the trained Loras works very well with different/specialized Krea-2 Turbo models too, as well as with combinations of other Loras.

    album3826Jul 21, 2026· 1 reaction

    @arkinson Actually got really good results! Thanks a lot for the guide and config! Next lora i will try to tweak the settings a bit to see what my pc can handle.

    arkinson
    Author
    Jul 21, 2026

    @album3826 Perfect! Please let me know your results. Especially whether you can handle 1024 resolution, more steps with lower learning rate and total training times.

    arkinson
    Author
    Jul 21, 2026

    @album3826 Btw.: I actually did some image generation tests with the Krea-2 raw int8 convrot model and the results look much worse than the Turbo Model outputs (just like plastic). Even without any additional Loras and 52 steps. Is this normal? 🙄

    I just did the suggested changes in my comfyui workflow: adding negative conditionning, cfg = 4.5 and steps = 52.

    [Edit: Ok, I got the trick. The raw model seems to need much more specific style descriptions. Just with the right prompts the outputs starts to become exellent 🙂]

    album3826Jul 21, 2026

    @arkinson I'm afraid I am the wrong person to ask, I didn't have much success with the raw model either, but when I tried it I used the turbo lora which lets you use 8 steps and cfg of 1. But great if you got it to work!

    arkinson
    Author
    Jul 21, 2026

    @album3826 I did just some quick-n-dirty tests only, cause I had not the time for more serious runs yet. Yeah, “brilliant” was a bit of a hasty conclusion, but very precise style descriptions. like in some Krea-2 samples, seems to have much influence to the output quality. Unfortunately I can`t test the full raw model with my hardware 😕

    album3826Jul 22, 2026

    @arkinson Back with results, I did 1400 steps at 0.00015 learning rate, 512, 768 and 1024 res, rank 32, and also disabled low vram mode on the advice of chatgpt lol. Took about 7 hours for 30 photos.

    Next time I might try 0.0001 learning rate, but keep steps and resolution the same since as I understood it too many steps can negatively affect the results (I might be wrong?)

    arkinson
    Author
    Jul 22, 2026

    @album3826 Wow, this sounds pretty good! 4 gb more vram and a faster gpu seems to make a difference. Never thought, that you could run 1024 resolution well with 16 gb vram. And it seems, 32 gb ram is not a big problem too.

    I did a short test with rank 32 on my system- but training time is becoming unbearably long....

    On your side I would try LR = 0,0001 with 3000 steps (that`s the suggested standard settings). You might safe every 250 steps for example.

    [Edidt: Arg - you are right: 1400 steps with Gradient Accumulation = 2 are equal to 2800 steps ].

    Btw. You can pause/stopp the running job at any time for testing the outputs. Resume will continue from the last step (or sometimes from the last saved checkpoint).

    album3826Jul 22, 2026

    @arkinson I haven't had time to test results a lot but yeah it looks really good so far!

    So maybe I should try either 1500 steps with gradient accumulation 2 or 3000 with 1?

    Thanks for all the help :)

    arkinson
    Author
    Jul 23, 2026· 1 reaction

    @album3826 Sorry for the confusion. Short anwer: I would use Gradiant Accumulation = 2.

    The general recommendation for relatively small image sets is 2 or higher.

    schschJul 21, 2026· 1 reaction
    CivitAI

    Is it normal to last days in the first stage (before even started the steps 0 of 750? Lol

    Caching latents to disk: 14%|#4 | 5/35 [16:51:00<105:48:12, 12696.40s/it]

    My system: RTX 3060 12gb+ 32gb RAM + 80gb SWAP FILE. Training pics: 35. I used your custom config unaltered.

    arkinson
    Author
    Jul 21, 2026· 1 reaction

    No, definitely not. Cashing latents is very fast. Cashing 40 hq (>2k) images stored on a hdd needs around a minute on my system). Anything suspicious: ram/vram/gpu/cpu/disk usage?? Damaged/ultra slow disk drive??

    Workflows
    Krea 2

    Details

    Downloads
    228
    Platform
    CivitAI
    Platform Status
    Available
    Created
    7/20/2026
    Updated
    7/26/2026
    Deleted
    -

    Files

    krea2RawLoraTrainingGuideWithRTX_v10.zip