MM H3 reference video ref2va lora
MM H3 lora trained on the reference to video model (ref2va), fun little project for me so I could make stupid videos. I'm using euler/beta. I would be cautious about mixing some other loras (test first) because when I try them they make videos look poor quality or cause smearing. I would also test without turbo first (I dont use them). use only ref2va trained loras. Don't forget with ref2va you can supply short 512x512 reference videos for your scene, that really helps a lot.
FYI: I dont use turbo lora, for various reasons, and I dont test loras with turbo loras enabled. I think it makes the distilled lora training related issue worse, rendering the lora useless.
Grok H3 Prompt Skill: https://grok.com/skill-link/6dc971af0dd0458b12ef523fa62696bf
How does a reference model work?
It doesnt work the way you are used to with traditional models. This is a reference model, it literally means you need to give it reference images because the lora and the model itself is trained on reference images. If you want a nude woman to suddenly run into the scene, give it a reference nude woman. If you want someone to look over and see two people fucking give it a reference image of what fucking looks like. Then let the lora animate it. The lora was trained on "take these reference images" -> "generate a video using them". it was NOT trained on text to video, it cannot generate a cock out of thin air, it needs a reference image to animate it. Otherwise use the fl2va model because it can be trained on pure captions. So if im doing a handjob scene, I might give it the main photo of the scene, a reference close up cock image, a reference oily tits and reference handjob technique photo. Doesnt mean I need all those, I could get away with one image but I cover my bases so the video turns out. The job of this lora (and what it was trained on) is to take images and animate it, it was not trained to create anatomy that doesnt exist in the references. It can sort of do it but it will look incorrect (deformed cock head or something because the angle of the cock is different than what it may have seen).
I put in the comments for each of my v1.2 videos how I created it. I uploaded a "what does a nude woman look like from all angles" photo to the gallery, reference it in your generations so the model doesnt have to invent anatomy as the camera is moving around. https://civarchive.com/images/140155060
you can even give this model video examples of how to fuck, or audio references of what it sounds like, the sky is the limit if you are creative. But its a bit of work compared to the simple fl2va model. this lora was trained on how to make reference images move, hopefully that helps.
You need GOOD prompting, use the grok prompt skill below unless you want to write 1000 word essays.
the two versions of this lora are described below, I would try 'sexytime' variant first.
Join our Discord: https://discord.com/invite/GCrr4Cj3D
WARNING: for the AfterMidnight ref2va lora you need to use euler sampler and beta scheduler, or you will get weird crap like audio artifacts or turbo tits. Listen to some of my videos and you can hear odd audio noise or see strange crap, thats from me using res_multistep. dont be a like me, cheers. this video uses the correct workflow https://civarchive.com/images/140029319 :)
Make sure the lora passes to your Basic Scheduler and Guider nodes or nothing will happen
Workflow: download video and open in comfyUI https://civarchive.com/images/140064227
Lora "Flavors"
there are 2 flavors of this lora (same dataset different training styles), one only one at at time:
sexytime: strength 1.0. Was trained for sex scenes and coherent motion, probably closer to what you want.
softer version: strength1.0, wasn't pushed on motion as much, focused more on detail. I run at strength 1.0. Was trained on detail and surreal style, think like crisp outputs and fantasy (flying dicks?)
Model: MM H3 ref2va bf16/int8
Training: mixed bag of 1500 videos of various sex acts, close ups etc
this is my first attempt at a ref2va mm h3 lora. Still working on my training pipeline for this one (I dont use ai-toolkit)
Description
FAQ
Comments (57)
Could you provide a bit more info on what this should be used for, specifically? Any sex act motion?
a mixed bag the most common sex acts
ty for your work cant wait to test.
no problem, I hope you enjoy it
thx for ref2v lora
my friend do a handjob lora for mm too hehe
already in some multi purpose ones. H3 is pretty clever. just describe it and it can pull it off.
@ROXBOT can you link me one? when i try the hand stands still
@akira I posted videos of handjobs for you in v1.2. If I can do it so can you. its a reference model, not a text to video model, give it a photo reference of what a handjob looks like if need be. the fl2va model doesnt need that but ref2va does. the lora is trained on reference images (take this image and animate it or add it to the scene), so my work flows have a reference cock, a reference doggystyle photo, a reference nude woman etc. YOu need this because thats what this model was trained on, otherwise use the fl2va first frame last frame model. the ref2va is more work for sure but you can do amazing things with it.
So for the handjob massage scene I had 1) reference massage scene 2) reference close up cock 3) reference oily tits 4) reference hand job hand placement. then I animate it making reference to those images. Did I need all of those, no I can do it with just a single image and it was 95% there but the cock head wasn't perfect so I just give it lots of images so it doesnt try to draw a cock from the base model (which has has crappy anatomy).
why rank 64 ?
seemed about right. just guessing
if dataset is varied, a bigger rank makes sense. Plus, this Lora doesn't specify a particular style.
@czlowiekbulwa888 its varied yes. think like 1500+ videos of almost anything
So far with limited testing. the best general purpose lora!!
thanks, im going to train it some more, but I think I got this model finally figured out
@sexgod1979 Nice to hear! And thanks for not paywalling it!
@Sanchez3 screw that im not pay walling any shit, im an open source sort of guy
work i2v?
the reference model is i2v
Basically. ref2v and i2v both use images to start. I2v uses the image as a start point, ref2v uses 1 or multiple images to be used in the scene. Think of it like a character lora without the lora. It is kinda diabolical how well it works in H3.
All samples seems to be in slow motion.
Is that the prompting or the model?
looks normal to me https://civitai.red/images/140075037
for instance I just did this one, speed is normal, its just how you prompt. https://civitai.red/images/140081100
@sexgod1979 Yes, those 2 have more movement, true. Thanks!
@EpicStuffs I suspect I can improve it though with training further, let me look at that
@sexgod1979 Awesome - thanks for your hard work!
slow motion demo?
looks normal to me https://civitai.red/images/140075037
I think its more a case of prompting, I can get normal speed videos or slow motion depending on the scene/prompt
for instance I just did this one, speed is normal, its just how you prompt. https://civitai.red/images/140081100
@sexgod1979 Did you use over 1.2 megapixels? It would cause to slow motion
@arthur021997 yes. slow motion from a large ref image?
@sexgod1979 megapixels mean the resolution output of your video, not ref image
@arthur021997 im generating at 0.8MP, some gens come out faster some slower. this very much appears to be a pacing issue in my lora training, ill work on it now and correct it. thanks for the heads up all
If you are running local and upscaling, fps mismatch can create a slomo upsale
Using the workflow from https://civitai.red/images/140064227, I'm not getting detailed vaginas. Are you starting out with vaginas already in the reference picture?
Yes, because its a reference model, its trained on "given these images + caption" generate "this video". Otherwise it would try to draw a pussy from the base model. Just give it a photo of a nude woman for reference, easy. It may generate an approximation of anatomy without references but its a light weight lora and has limited capacity to retrain the entire model on that.
for pure t2v generation without references I would use the fl2va model not the ref2va model as most of us provide ref images, and my lora training does the same "given these references + this caption generate a video". The best solution would probably be a full fine tune of the ref2va model to teach is anatomy, but its a big job
@sexgod1979 "It's a lightweight lora" lol.... no it isn't! It is rank 64 and 1.11 GB. There is nothing lightweight about it. Nice work though!
@kermitfrog1202 well ok, yeah youre right. its not a tiny lora.
Use a detailed reference pic for those bro.
Where do we get qwen3vl_32b_minimax_h3_int8_convrot.safetensors which is part of your workflow?
All you have to do to find almost any model is paste the exact name into google and click the first HuggingFace link that pops up
@jayhartford there is no hugging face link that pops up for this name though... I have tried this already
@makemoneyrajp971 bruh…just google it
@makemoneyrajp971 Come on, literally the first Google search result is a link to the model on HuggingFace...
do you have any better penis LoRAs?
only in my pants. kidding, no sorry not at the moment.
also since ref2va model is conditioned on references, the easy solution is just add a penis reference close up to one of the reference image slots and make reference to it in the prompt then the model will draw it perfectly
if you are doing ref2v you can reference one or your own. lol
@ROXBOT of course, I tried that though and I got this error message in comfyUI that said "Sorry dick too tiny, please press any key to continue"
@sexgod1979 That's what photoshop is for bruh lol
@ROXBOT <cock is now 19 inches long>
@sexgod1979 I tried this but it is just being ignored. And the photo was good and prompt as well. It takes into consideration other pictures but ignores the penis one.
@pauliess949 It works. My guess is you have your lora wired wrong, youre using the wrong base model or some aspect of your prompt/workflow is incorrect. Since I cannot see your workflow its difficult to know which. Use my modified workflow or show yours. for the record I use unpruned int8, ive never used or tested pruned. Im sure it wouldnt make a difference but I dont know. This lora was trained on unpruned model. also I would make sure theres no turbo or speed up lora as I dont know how that would effect the lora (I dont use them either because I dont like the results)
@sexgod1979 i will try, thx.
@plss949 I was having this problem too at first with a close cropped penis image. It copied it correctly when i used a different reference that included the penis and part of the body.