This is a port of my dispatch dataset from LTX to H3. It's trained on the fl2va model. All samples are t2v on 20 steps.
Right now only 3 characters are trained, but I want to add in all of them. I wanted to do this so I could see if H3 can do multiple characters in 1 lora like Wan/LTX does. So far the answer is yes it can, but the voices are strange. Consider this like a proof of concept beta version.
The voices are trained, along with music from the game. But when you have multiple characters in 1 scene speaking, their voices tend to mix. Especially invisi and blonde blazer in the same scene.
The voices also are like 90% there, but still not super accurate. I also have seen this same problem when using the ref base model. It doesn't like to have multiple of the same gender voice in the same scene. I will keep training and tinkering to see if I can find a solution or not.
Style Trigger:
"Stylized 3D cel-shaded comic style."
Invisigal:
char_invisi, a woman with short dark hair accented by a purple streak, wearing a purple jacket over a dark top and distressed black jeans
Blonde Blazer:
char_bb is a woman with long blonde hair and a blue mask, wearing a blue and yellow superhero suit with a yellow cape and a red diamond-shaped gem on her chest.
(no costume, powerless)
char_bb has long, wavy dark brown hair and wears a strapless evening dress with long dark blue opera gloves
Robert Robertson:
char_rr has short brown hair and wears a light blue button-down shirt with the sleeves rolled up to his elbows and dark trousers with a beard with stubble.
Description
FAQ
Comments (8)
Please make a dragon ball Z style H3 lora...
H3 already knows dragon ball
I'm gonna keep training and testing but unless I can figure out how to fix the audio issue with multiple characters, I might not add anymore characters. Since it would be a waste if they dont have the voice.
I've had some success prompting multiple voices of the same gender before in Ref2VA mode for other prompts, though there was some confusion sometimes. Maybe there's a way to sort this out so that if multiple characters in the lora are speaking. Are your examples FL2Va? I'm going to run some experiments running the model, see what I can get from it. Could you maybe upload a segment of your dataset?
Hmm done some experiments and hearing what you're hearing. When you're training is it one dataset you're using or multiple ones?
Is it possible to train on just audio files? Maybe they need some extra specific data to push it a bit more. Some clean source audio of spoken lines without anything else in it.
Is this a style lora or characters lora ???
both