09-20-2026 - Add Starfire Lora with proper audio
I'm still trying to learn the right way to clip videos and caption for good audio training. My plan is to make Starfire and then the next release will be Starfire and Raven in the same lora. If I can, then I'll keep adding characters with each release.
First Lora style works great but has no audio training.
09-17-2026 - Initial release
I tried to train these characters into the style lora:
Starfire - strfr, red hair, long hair, green eyes, orange skin, crop top, midriff, miniskirt, thighhighs, thigh boots,
Raven - rvn, purple hair, short hair, purple eyes, grey skin, forehead jewel,
Jinx - jnx, jcolored skin, grey skin, hair horns, pink eyes, pink hair, choker, long sleeves, pantyhose, striped clothes, striped pantyhose, black dress, wide sleeves, collarbone,
Robin - rbn, black hair, red shirt, green sleeves, short sleeves, mask, superhero, domino mask,
Description
TnC1tr0n,
FAQ
Comments (40)
Nice to see you making H3 Loras!
Thanks! I'm surprised at how well training is working with just static images.
If I can figure out how to train and prompt them locally on a RTX 3090 then I might train H3 full time.
@CitronLegacy I second that idea for you lol
You got upgrade :D
Maybe try Fizgig?
@RisingV This https://github.com/shootthesound/Fizgig?
It seems interesting! I'll have to see if I can set it up!
It looks like it can train Krea 2, as well? I originally rejected Krea2 because it looked like I couldnt train locally but if I can it would be nice to try.
@CitronLegacy Yes, I successfully trained Krea 2 LoRA with it: https://civitai.red/models/1890171/kyhu-artist-style-k-y-h-u-old-pen-name-of-iafhy-aniilkr2zit?modelVersionId=3214341
Though it was awfully slow on my pc, when I tried it cry
Could not use int8convrot training since it didn't support offloading and fp8 is not supported natively on rtx 30xx so it used fp8 weights but did computation in bf16^^
@RisingV Ok cool. I might try Krea out but honestly Anima and IL fill the image gen role perfectly.
Slow and steady works! Better than ZIT training which had a 30% chance of crashing my computer lol.
Its so annoying when training only works on 30xx or 50xx!
@CitronLegacy Yeah, I haven't getting a full grasp of Krea yet, but it has an advanced text encoder compared to Anima. Thanks for buzzing my model!
thanks for lora!
Glad you like it! :D
Thanks for making this, pretty good for a first attempt at it, though current H3 training does make it hard to prompt multiple distinct characters and prevent bleed. Hopefully there's some info out there that would help if you decide for Version 2. would be amazing if you could train the voices as well.
Thanks again!
Thanks glad you like it!
Yes my research (and Claude Sonnet's research) suggested that its not possible to train multiple characters into a style lora, but I wanted to try anyway!
It worked much better on this lora https://civitai.red/models/2943842/kim-possible-series-style-minimax-h3
Biggest problem with training voices is I dont know how to get the video clips I need to train the lora. Once I get good video/audio data sources I'll be able to make really good things.
Thanks for replying! @CitronLegacy From what I've heard, the hardest part is the fact that training two women or two men cause the greatest clash. I'll check out the KP one as well and Kim and Ron look to be working well.
For voices, how many video clips do you need? There's a number of clips on YT you could potentially extract from and then separate music from. I've got some clips too from when I've experimented with TTS using their voices.
@Rosettasees I guess if I was going to focus so much on audio for specific characters it would be easier to make a character lora rather than a style lora. That way it can focus on learning one voice the correct way.
It looks like the internet is saying 50-200 clips to train with videos but maybe we can get away with 10 video clips. I guess if there was a really good set of 10 clips for a character it would be a good first start at training a character lora with accurate voice.
It would be super fun to have Starfire, Raven, or Shego's voice working with a character lora.
@CitronLegacy Hmm, well for Starfire this video might be a place to start for material?
https://www.youtube.com/watch?v=gWN0sejD7Fg
@Rosettasees Thanks for the video!
If CitronLegacy thinks that just dropping a massively nostalgic, retro style like Teen Titans for Minimax H3 will make us download it instantly? Well... that's correct, thank you! DBZ or Gundam next, pl0x tnk u🫶
LOL Glad you like it!
I have two DBZ datasets but I'm not sure about the quality so I havent tried training one yet. The biggest concern in that the AI captioning probably wont understand the difference between the characters and it will mess up the training. I'm still learning so we'll see if I can figure it out.
@CitronLegacy I've noticed just from talking to LoRA trainers that Krea2 and Anima both have better character recognition than older models . And while people are still learning and working with Minimax, I'm hearing similar. You make good stuff and are clearly smart and skilled. I have zero doubt you'll figure it out sooner than you think, and better than you'd guess. I'm looking forward to it.
@7117 Here is a Dragon Ball Super lora https://civitai.red/models/2945909
I really love the Super Broly movie so I trained on the action clips and I think the lora makes some cool videos.
Thanks for the compliments! They were good motivation to try making this lora.
@CitronLegacy Fantastically done, this is going to make for neat action scenes. I sometimes use AI to create fun stories or videos for my nieces and nephews, so I'm pumped for version 2 as well that includes non-action segments. Great choice, you and I both know that DBS: Broly is arguably the best modern animation style they've done, so you picked a fantastic platform to give yourself options! Thanks for these, as always
So cool, any recommendations for training settings and captioning?
Thanks! Honestly I'm surprised it worked so well.
-- Copy pasta from another comment (I need to post an article for this so we can all collaborate in one spot lol) --
I'm training only on images so far. I'm using the same datasets from my Illustrious loras but I'm using captions instead of tags.
I made 2 loras that might have failed but 8 successes out of 10 different ideas is pretty good.
I'm targeting 100-200 images.
However smaller datasets seem fine
34 images - https://civitai.red/models/2943813/minimalistic-style-minimax-h3
26 images - https://civitai.red/models/2405530/civchan-yandere-buzz-queen-anima-and-h3
39 images - https://civitai.red/models/2943858/perona-ghost-princess-or-one-piece-minimax-h3
I'm using only the defaults on the CivitAI trainer. I usually use a different tool for captioning but I used the CivitAI caption-er for https://civitai.red/models/2405530/civchan-yandere-buzz-queen-anima-and-h3
I'm still learning and a have a few ideas on how to caption the datasets better but it works right now.
@CitronLegacy Thank you! I am just preparing a dataset to train a style myself (that minimax can't do well it seems based on very limited testing). I think I'll try to use 50 images and see how it works...
The visuals are spot on! but your recent Minimax H3 loras seem to have used video/audio data that had dialog, but wasn't captioned for it, because they see nonsense dialog as part of the standard output.
Thank you!
Yes unfortunately I'm not training the ideal way. I'm training with images only so there is no audio in the dataset. I was hoping that good prompting for the audio would produce good audio.
As for training with audio I'm still trying to wrap my head around how to make a good dataset of ideal clips.
I don't have the source video files to cut clips from and the video quality on Youtube seems like it would damage the video generation.
I think it would be super fun to train the correct way with proper audio-video clips but I dont know how to get the data.
@CitronLegacy Okay. Maybe I can help out.
@Jellai If you can find a source for the clips please let me know! It would have been nice if there was a site like this for video clips https://fancaps.net/tv/showimages.php?37768-Teen_Titans_Season_1
@CitronLegacy I messaged you
@CitronLegacy There literally is. It's called pitatebay. Then just use a video editor app to cut it into clips and save the individual clips.
I always loved this Teen Titans, thank you so much for sharing this! Are you planning on adding more characters? Thanks!
My pleasure! Yes I'm planning starting over with just Starfire so that I can work on getting her voice correct and then I'll start loras that support multiple characters and their voices.
@CitronLegacy That's awesome!!! Can't wait!! If I may Robin and Aqualad are my faves! I hope they are included if you're able! :D
@VioletCrow Yes for sure Robin. Daimon is my favorite Robin but the one from this Teen Titan's is my second favorite Robin.
Not sure about Aqualad. Dataset creation for H3 is pretty time consuming because I'm so new to it.
@CitronLegacy No worries at all! I don't know the first thing about model training, but I can understand it is very hard. I would just be happy if robin was there lol. Yeah this show is so nostalgic! BY THE WAY! I just noticed you uploaded a stirefire one its INCREDIBLE!! GREAT JOB!! :O
Azarath metrione ZINTHOS!!! Thanks !
Hearing that line makes me want to making a Lora for Raven's voice! Thanks for the comment!
Pretty well done Starfire, her voice has definitely been captured properly and animation looks good, though color grading is bit off, though that could be for any reason. Looking forward to whatever is done next!
Thank you for the feedback! I'm really glad to hear that it is good.
Yes I agree the colors are weird. It might be because I had some dark scenes that confused the learning (I tagged them dark scene but apparently that didn't help lol).
Raven & Starfire in the same lora is planned to be next in this series! (Hopefully its possible to train two different voices without them blending together)
Prompting the lighting is always a challenge, maybe try to prompt the light scenes as well so it can differentiate them easier? Like 'well lit' or 'dimly lit'?
Looking forward to what the Starfire & Raven result is!