This lora works on i2v and ref2v INT8 Convrot
training time [2 hours 9 minutes]
rank16
resolution 512 (because 64GB of RAM and a 5090 isn't enough)
[Guide]
Captions were done using SDXL. Adding more words destroys the lora because many words can mean the same thing and different things in different contexts. This applies to all models. I found this out when using Krea2 by using the krusty krab restaurant and krusty the clown (New benchmark???). Minimax uses the same family so it shouldn't be different.
If you train yourself eating an apple, you can just say, eating ohwx. Anything more causes conflicts within the lora. The model already knows what this eating apple action is, buts it's not trained in detail. You can now say a person eating an apple with and without using the trigger word.
When you train a lora, you are training it as if your model was starting from zero. Reason being is the lora will weigh more than the model knowledge, your lora becomes the #1 priority over the model knowledge. For example, if you want a GTA style lora, you need images that depict it as the entire dataset of the model, meaning 1000 of thousands of images, even millions. While the model can guess it with around 20 images, it's better if the image exists of what you have in mind inside the dataset. Any guesswork will increase the chance to bleed concepts.
I've had prodigy shit the bed after 1200steps when I did a krea2 lora, somehow the dataset I used didn't like it, but it works fantastic for most things I do.
I use caveman talk since it works and is more efficient on tokens, saying "once apon a time" is just a waste of tokens.
The model doesn't know what a cowgirl position is in great detail but if you feed it an existing image it knows what to do without this lora. What i'm doing is giving the model CPR to revive the cowgirl position meaning. By captioning the different angles, to not taint the knowledge of the other words, I keep it short and simple. This concept would be much better to do a full model training with longer descriptions.
If you gut out words in your natural language prompt when your generating images, it will still work.
If a pregnant woman says she is pregnant, does she describe in great detail how she got pregnant? Mostly likely not, you can assume that penis goes in vagina and that's that. If I were to break down what position it was, most likely doggie, missionary or cowgirl.
When you caption, you use the words inside the model, for example if you say, a photorealistic image of an anthropomorphic mouse, when the image is actually 2D. It can turn that image into photorealistic, which is basically reverse psychology. Hence why words destroy the lora, because one can say hyper-realistic, surreal, photorealistic, realism, they all mean different things in context of the model since the model is trained on what the dataset was captioned with those words. Photorealistic isn't actually photorealistic and produces more of a 3d effect. But if you remove the photorealistic tag for it, it fixes it, but now you can't rely on the 2d part of the dataset because now the photorealistic tag contaminates the other tags that determines the style of the lora, the other styles will still work, not so well.
Qwen is such a mess, I never had issues with Gemma4, so I could prompt it normally without much issue. They should have used Gemma4 instead. LTX2.3 prompt adherence was amazing but LTX 2.3 dataset is lack luster and not able to get the full advantage of Gemma4's abilities.
BTW, captioning is also RNG, using prompts inside your AI generator also causes RNG combined with the lora captioning. So your captioning strategies can be RNG, however, the way I captioned it is consistent between runs and different concepts and models. However I haven't tested to see if every caption actually works, because that is also RNG, the main one does work and i've seen it get creative which actually makes the i2v portion more interesting.
Sometimes you don't even need to say anything more than cowgirl position for it to work, thrusting her hips or thrusting her body is enough.
BE WARY OF THE MYSTERIOUS HAUNTED NIPPLE LIGHT AND THE EVIL LIGHT SWITCH
Don't worry about the dragon, she's just casting a curse on you in her native language
cowgirl position
cowgirl position from the front
cowgirl position from behind
cowgirl position from the front left
cowgirl position from the left
cowgirl position from the right
moving hips while in the cowgirl position
moving hips while thrusting in the cowgirl position
inserting penis into vagina with her own hand
cowgirl position while squatting
thrusting slow
male moaning
male pov
male pov low-angle
male pov mid-angle
male pov high-angle
male pov low-angle to mid-angle
male pov low-angle to high-angle
male pov high-angle to low-angle
male pov high-angle zoom-in
closeup male pov low-angle
female pov
female moan
female moaning
female soft moaning
mid-angle
high-angle
low-angle
low angle focused
wet pussy noise
smacking noise
moaning together
bed creaking
a woman is climbing onto the man laying on his back to get into the cowgirl position
kissing
after sex
orgasm together
female orgasm
female muffled moaning
wet nosie (I miss spelled it, should be the same as wet pussy noise)
3rd-person
wet sounds
subtle female moaning
male groaning
closeup of breasts
female quick moan
female grunting
male exhale
thrusting her hips deeply
female orgasm whimper noise
sucking on breast
+more
Description
I hope I don't get banned for this,
**Please read the license agreement for MinimaxH3 before using this Lora**
FAQ
Comments (13)
This looks like you invest a good amount of time into this lora. I have a question. I was trying more than intensively to create like POV standing carry videos with Wan 2.2 where from point of view holding a woman by her waist/hips while standing and she also holding on his hands and see the bottom thrusting and the face at same time. Well I never fully acomplished this closest I ahve been was cowgirl with the hand composition but never standing carry POV as I wanted.
So with Minimax H3 seems have better understanding of prompt text. Do you think that your lora can help with a standing carry POV or it would not work and I would be back at cowgirl or some kind of it. Or do you have any appetite to take this task, if it is ever possible to do a lora that would help or do you have tips that could help with that?
Something like this would require a 45k computer, or rent from runpod for like 15 dollars for better results. I don't have the hardware to do it properly, but it should work in theory. Minimax is only good because of the dataset, not the text encoder. If you want to do something like this, you need more than 3 seconds of video in your training data. Using my lora will help guide it better and should transition well, since I spent everyday learning how this model works since its released. I also heavily modified the dataset which helped a ton, since I don't have enough RAM and VRAM. If you do decide to make it, just make sure it only sticks to that concept and nothing else. I could do it but I don't feel like doing it since (Will take about 15GB of my storage, takes time off of future gen and training, Minimaxh3 takes 45 minutes to generate 1 video on a 5090 with 64GB RAM and it's not even guaranteed to be good. I use higher than normal settings to get better quality when I need it. Lowest I've seen was 2 minutes and 34 seconds. Most of the time it's 5 to 15 minutes depending on length. You can follow the prompt table below to see how I prompt it.
@brand175 Thanks for reply. Minimax H3 is interesting, one interesting thing I got was when I prompt someone talking and there are more people in the generation and does not really have precise task to do, that the 2 or three different words can affect the, have should I say it, the mood and the overall generation to some degree. My testing was not focused on that but it was big enough to notice.
Yeah, about the training with the system of Minimax H3, as you mentioned with one 3 second video, I could do reference to that video and create bigger dataset or edit the generations to get better result. Currently I trying to get similar results of generation to WAN 2.2 but it seems that I am not fully doing the thing or I had the WAN 2.2 workflow tuned enough to have a better quality and consistency of character in Wan than I have in actual testing with Minimax H3. But new fine-tuned models will come up, Dasiwa made a hybrid I must test if the result goes up.
Again thank you for your reply and the lora!
How does this go with anal? or does it do the slip into vaginal thing?
brand175 doesn't like anal, so its not trained on it. But you should be able to use this lora at lower strength with an anal lora.
Thx for your wall of text. I am experimenting with musubi-tuner dev with minimax support. And yes short captioning is the best. The model knows and understands a lot. Many creators on this site Caption like that : Her skin highlighted by rays of the moon, at the night. The moonlight plays on her skin, as she playfully at intimate angle of view..... insert .... Please stop.
Nice, finally a good "mounting" from the top, not a weird sliding in!
It does slide in when you ask for cowgirl position if you don't caption the natural movements before insertion. You can enter in at different stages because I staged the dataset to be like that, AI doesn't like it when you state everything that happend in that time in lora training. So I broke the insertion section into three parts (Insert, thrusting, pulling out). This is this more efficient, and it uses less VRAM. Not only that, AI can't know to start from said caption in the middle of a 10 second video from the dataset. The first frame needs to start with the action and end with said action, per section. It's like staging a rocket in KSP, using only a booster can get you so far.
@brand175 I understand none of your words but thank you, i was simply saying your lora works very well :D
Also works well for t2v. If you could make a general vagina Lora that would be awesome! This one has the best vaginas I have seen for H3 by far
Absolutely amazing lora!