Example prompt:
missionary sex, POV overhead view of a woman lying on her back with her legs spread having sex with a man. She's on the couch. She is moaning with pleasure. A man is thrusting his penis back and forth inside her pussy rapidly at the bottom of the screen. The camera is zoomed out and holding steady. Her auburn hair is long and curly.
Version 1.2 Updates:
I used the same training videos as before, but this time I blurred the faces. I hope this created a non face altering version, and it definitely does seem to work better with character loras, and the facial variety also seems a bit better.
I used https://github.com/ORB-HD/deface to blur the faces in the training videos, and then added "woman with a censored and blurred face" to the captions.
Version 1.1 Updates:
I added a few new videos to train on, zoomed in more on the movement itself to hopefully train that a bit better. I also lowered the learning rate to 5e-5 and bumped up the number of repeats to 30.
The result seems to work much better at lower strengths and hopefully better with character loras now.
Note that this training took 8 hours compared to v1.0's 1.5 hours. There's probably some sweet spot for learning rate and training time to get good results, but it would take more experimentation to figure it out.
I don't feel comfortable sharing the exact mp4s I trained on, as they were just ripped from online sites and I don't really have the rights to distribute them. However, I will include my training data for the config files and captions so that other people can more easily get into training. I was surprised at how quick and easy it was.
I included an example workflow in the training data download (I can't find a better way to upload the workflow), which shows how to make it work nicely with multiple LoRAs, and has dynamic prompt support.
v 1.1:
I trained on a 3090 using 11 3 second videos (24 FPS, at least 50 frames each) and it took around 8 hours to do 20 epochs with 30 repeats.
v 1.0:
I trained on a 3090 using 8 3 second videos (24 FPS, at least 50 frames each) and it took around an hour and a half to do 20 epochs with 10 repeats.
Description
Trained with a few more images, some zoomed in on the movement itself, and at a lower learning rate.
It now seems to work better at lower strengths and is probably better with character loras.
FAQ
Comments (36)
Any chance you can do one for anal? There aren't any loras for it yet :(
V2 is a clear step forward. The movement feels much more natural now ;)
What a time to be alive!
How much VRAM it needs btw?
hunyuan vid? 12gb for the fp8 model and default encoder but you can find a smaller gguf and text enconders than could probably work with less, but you still need over 24gb of ram
@ginx Im running at 16gb ram, works slower but it works
Has anyone tried making these in stereo (cross eyed) yet?
are you planning to update this? What I would like to see is to include different angles. You nailed the center front view. Now seeing it from different angles would be a nice upgrade!
Also the Hunyuan model seems to have such broad knowlegde of poses and movements that it would produce amazing results. I'm trying it with the laying cumshot lora right now and depending on the location it produces very interesting scenes tht work very well. Man, I'm so used of problematic finetuning of Flux and SD3 that I'm impressed every time how well this is just working =D
I'd imagine if I were to try different angles, I'd make it a whole new Lora, as this one is kind of at the limit of how long I'm willing to wait for the training on one thing. I have no immediate plans though, so maybe someone else will do that first.
Hey, Great Lora! May I ask what soft you used for training the lora: Diffusion-Pipe or Kohya-ss Musubi-Tuner? Do you think 3090 and 32 ram is enough for training from videos?
I used diffusion-pipe. I'm not sure about the amount of RAM needed, I have 48. 32 would hopefully be enough, but I haven't checked to see how much it's actually taking. The 3090 is what I use.
@dtwr434 thank you, kind sir :)
@dtwr434 You've got a decent rig there, mate! 48GB VRAM what a flex!
@YesPleaseProceed nah, he says it's ram, his 3090 has 24 gb vram i assume
I have an interesting observation. I have been generating at lower resolutions to save time. The 1.0 worked fine, but 1.1 was all broken. However, I mixed the two with 1.0 at full strength and 1.1 at ~0.4. The results are great!
Can you share with in mixed model?
Hey! It's back. Dunno why this model was gone for a day or two.
Where has good guide how to train models?
Hey, i think it was removed because of the hunyuan license. Maybe you need to censor the showcase videos. The nsfw lora was also removed https://huggingface.co/TheYuriLover/HunyuanVideo_nfsw_lora
but its back on huggingface without images.
Your lora is great! , can I have a footjob with lora, thank you very much!
Just noticed you put the dataset.toml in the zip, thanks, gotta try to copy your setting on my next training and see if I finally get somethhong 'moving' :p
EDIT : I see so just [1,50] in frame_buckets, the interesting part was 30 num repeat, (I only had 5, probably gonna try more in my next run and see where it goes) now one more question remain for me, does having video training clips in a dedicated video folder (with corresponding path in the .toml file) is needed or putting all in the a default images folder (images plus videos alike) is the same deal ? because I put everything in the same folder since the start
For the repeats, it's just a question of whether you want more repeats or more epochs. Since I'm saving off the model every 5 epochs, I opted to go with more repeats and less epochs so it didn't take up so much disk space. The number of steps is the more important part, I think. I'm usually ending up doing <number of videos> * 100 for the total steps, so maybe just pick some combination of repeats/epochs that gives you something around there.
I think it's up to you how you structure your directories. I kind of like having videos and images in separate directories so I can be more explicit about what resolutions to use for each, but I think as long as you specify the list of relevant resolutions at the top of the file, it should figure it out for you.
@dtwr434 okay thanks for the explanation, it make sense
@dtwr434 And LR, I see you've changed that!
@azeli Yeah, I still haven't totally figured out how much of a difference it makes, and didn't want to wait for training as long this time. I'm not sure I can tell much of a difference in the movement with this LR from the last one.
@dtwr434 I'm only as far as trying with images, and keep ending up with major artifacts in the background. Just doing an extra 300 steps for a total of 1300 so will check that but otherwise I'm stumped.
Captioned using florence2 and added a trigger word at the front like I did with Flux, but maybe I'll try with no captions at all.
How many steps are you doing, though I assume video is completely different. HAve 2 x 5090s coming very soon can't wait as will make training probably 10x faster than my current setup
@azeli Hmm, based on nothing else, I'm guessing the issue is with the images themselves, though I don't know why it would only affect the background. Are you using stills from video, or high quality photographs? Screenshots from video might have artifacts in them. When I train on images, I seem to get better results just sticking with 512 resolution instead of 1024, though it depends on what you're trying to do, I suppose.
If you have something like 20 images, I think around 1000 steps should be fine. You may need more or less steps depending on if you have more or less images.
@dtwr434 I'm using 1024 images, fairly high res nothing wonky. Exactly 20 images as well.
I may have been a bit lazy with my prompting for tests, I've prompted much more movement and they are coming out far better but still not like some character loras here.
One question though, do you train on the fp8 or bf16 model, and then what do you run the loras on? I always generate on bf16, but the training example in diffusion pipe is fp8 so thats what i was using
@azeli I use the fp8 safetensors file for both, though I have "base_precision" set to bf16 for inference, so I'm not really sure how that works. If you're out of other things to try, I would definitely give 512 resolution a try if this is just for a character lora. You don't have to resize the images or anything, just specify the resolution as 512 in the config file.
@dtwr434 I was wondering about that but the readme isn't clear about resizing/cropping. I changed bucketing to false and set aspect ratio to 1,1 etc as all my images are 1024. It definitely resizes them?
What I'm finding is that any prompt with full body in it such as walking or standing, I just end up with a frozen frame with no movement... FAIL :( :(
@azeli I'm pretty sure it does resize them, yeah. If you look at the cache folder it creates, you'll se references to the resolutions you're actually using. No clue about the lack of movement though. The problem I ran into with 1024 was that the likeness was not anywhere near as good. I wonder if it's a problem with the captions or something.
How is the LORA, with diffusion pipe or other? formed, what are the necessary PC capacities?
Where is the workflow json file in the training data download? Was it removed in an update?
Yeah, possibly. If you switch to the 1.0 and download the training data there, it should be there.