WIP pretty bad but better than nothing. The example is text to video. Maybe it helps with image to video but I have not tested.
v0.2: starting to look better but I did not test it much, maybe it is pretty overfit
update: with more steps of training it broke down. Looks like we need the undistilled base model or a training adapter first.
update 2: Don't use, was an early experiment, there are much better lora now
Description
FAQ
Comments (58)
was this trained on images or videos?
videos
Damn, i hope this will get a lot better :D
But thanks for your efforts <3
me, too. Text to image is not great, yet. But maybe it will do well with image to video but I can not test because my pc is training the lora :D
@herpderpmerp i will do some tests later. Do you use it on str 1?
I can get decent T2V vaginas, with detailed AI assisted prompting on the regular int8 version without any loras, basically detailed describing and using anatomical accurate language, not sure if it's doable on the pruned version though. An anus however, that is very difficult to prompt for well. But it varies shot by shot, on how to best handle the wording.
@Lora_Addict yes str 1. step 1000k just finished already much better. uploaded it.
@PahviKahvi Mind sharing your prompts? That might help a lot of people. Thanks 🙌
@GlowingGuardianGirl I just described what happens in my words, nothing special. But I only use 15 vids so I hand made the caption. something like "a woman with blond hair, she is naked, her vagina is visible, she spreads her vagina using her fingers"
@GlowingGuardianGirl Well, admittedly the stuff I generate tends to be very close-up, so I haven't really tried to get it working on wider shots, but here's a shrinking fetish prompt, where a couple young women insert the viewer inside them. Works... decently for me, 0.4res 16:9, 10s. (I also have sol-attn installed, and using it with 0 start and 1 end positioin, so it changes final output quite a lot, but looks fine here without it as well. Still far from good enough, but it at least doesn't look like someone cut a wound there with a knife. It's a long prompt, ai generated, might get better results with something more concise (again this is very close-up shrinking fetishy, so if you don't want to see that, don't generate with this. seed used is 1024988996462166. Edit: Oh right I also have a uncensored heretic text encoder in use, instead of the defautlt one, not sure how much it changes though.
Prompt:
integrated_multimodal_description:
[Shot 1] Authentic live-action filmed footage from the strict first-person viewpoint of a person one millimetre tall, already pressed between an enormous smooth fingertip and the glistening spread vulva of an incredibly beautiful, sexy and attractive 19-year-old very young adult woman (S1: Miia) during humid summer evening lamplight. No establishing shot, no cut, no pull-back; the frame remains sealed to macro surfaces throughout. Mixed warm lamp light and dim golden evening light, shallow depth of field, fine organic film grain, slight natural handheld shake from the tiny body trembling. Miia lies propped up on her elbows, her upper body angled toward the viewpoint, head tilted forward, looking down between her smooth bare thighs with amused fascination. Her vulva fills the entire frame edge-to-edge as an ordinary full-sized young woman's correctly proportioned vulva: plump outer labia lying open as smooth rounded pillows; beneath them, delicate inner labia parted slightly by the huge fingertip, exposing a soft wet crease rather than any open cavity. Near the top of the frame, a soft clitoral hood forms a delicate layered sheath covering the glans entirely; its prepuce continues as a faint smooth ridge that disappears down between the inner labia — no exposed glans, no raised button. The vaginal entrance lies low in the frame, a soft tilted oval ring of moist tissue framed by the inner labia, far below the hood. At the extreme lower frame corners, the viewpoint's microscopic male hands, each under 0.1 millimetres long, grip the fingertip's friction ridge, fingers curling and re-gripping. Low against the upper frame, Miia's pale-blue fitted scoop-neck T-shirt rucks up beneath her full natural D-cup breasts, leaving her sun-warmed midriff bare; her soft oval face, rounded cheeks, grey-green eyes and light-brown half-up waves look down with mocking delight. At the upper right edge, Sanni's face — sleek below-shoulder dark-brown hair, warm dark-brown eyes, faint nose freckles, light-olive skin — leans in, grinning, and she lets out a bright feminine gleeful giggle,
<d>[Giggle, gleeful, high-pitched]</d>
then speaks with smug mocking delight,
<d>[English, smug, gleeful, mocking] "Fuck, I want him inside me."</d>
Miia releases a low breathy feminine moan from above,
<d>[Moan, low, breathy, pleasure-filled]</d>
From 00:02.500 to 00:06.500, Miia drags her fingertip slowly downward through the wet labial folds. The viewpoint follows, yawing and pitching with the motion; the inner labia part around the lens as smooth glistening sheets, a long wet adhesive squelch trailing the fingertip. A thick glossy strand of fluid stretches across the frame, wider than the viewer's whole body, then snaps with a sticky pop; a single droplet detaches, wobbles, and rolls downward past the lens, larger than the viewer's head. The tiny hands scramble, hooking around a microscopic fold of skin, then splaying as the fingertip changes direction, palms sliding across the friction ridge. Sanni's delighted giggle comes from above,
<d>[Giggle, soft, delighted]</d>
while Miia's breathing turns ragged,
<d>[Pant, breathless, excited]</d>
and she lets out a sustained rising feminine moan as the fingertip smears down to the vaginal entrance,
<d>[Moan, sustained, rising, pleasure-filled]</d>
At 00:06.500, the fingertip presses the shivering viewpoint against the soft oval ring at the bottom of the frame, warm fluid seeping around the pad. The viewpoint turns its head sharply upward toward Miia's grinning face; she looks straight down through the gap between her breasts and speaks in a breathless teasing voice,
<d>[English, breathless, teasing] "Just a little toy for us now..."</d>
Sanni gives a stifled laugh from the frame edge,
<d>[Laugh, stifled, amused]</d>
and Miia exhales a sharp feminine gasp as she presses the fingertip forward into the entrance,
<d>[Gasp, sharp, aroused]</d>
At 00:08.500, the fingertip pushes the viewpoint fully through the vaginal entrance; the outside light narrows to a thin warm line and vanishes as the labial folds close behind with a wet squelch, sealing the viewpoint inside. The final frame is complete warm near-darkness: smooth, pliant, moist mucosal tissue pressing in from every side, thick syrupy fluid coating the lens, and the microscopic male palms flattened against the contracting tissue. Muffled feminine reactions arrive through the flesh:
<d>[Moan, muffled, body-conducted, pleasure-filled]</d>
<d>[Giggle, muffled, delighted]</d>
overall_soundscape:
Warm summer-evening ambience outside the room — distant birdsong, faint lake lapping, a breeze through birch leaves — muted by wooden walls. Close at the lens: the tiny body's rapid heartbeat, short breaths, and the soft skittering of microscopic hands over the fingertip ridge. Miia's torso, propped up and leaning forward, gives her breathing a clear resonant quality from above; her abdomen rises and falls visibly in the upper frame as she pants. Her wet labial folds squelch loudly around the moving fingertip, fluid strings snap with sticky pops, and the detached droplet lands with a thick wet impact. Sanni's bright giggles and stifled laughter arrive as close feminine sounds with a slight upward pitch from delight. As Miia speaks, her breathless voice comes from above, slightly enclosed by the thighs framing the viewpoint yet still clear. During the final insertion, external air sound is replaced step by step by warm internal wetness: rhythmic low pulses from the surrounding tissue, dense mucosal squelches, a damp suction seal, fluid shifting in slow viscous waves, and muffled feminine moans and laughter conducted through flesh. Every impact vibrates through the enclosing tissue and into the lens.
non_diegetic_music:
No music. Diegetic audio only.
@PahviKahvi That's what the model needs to get good results I guess. Thanks a lot. What do you use to get your prompt, a local ollama llm? Cheers 🙌
keep going
gotta start somewhere right? give it a week and one of the nerds will find a way to make this easy 🤞
Horniness is the primary productive force.
@1824649533931 uga booga agrees
Tip: you can literally use the ref2va model, give it an image of any **** you like as reference and it will snap it anywhere you ask it to perfectly without lora.
Yeah but i2v isn't as fun as t2v
@lifania REF2V isn't I2V, and you might be underestimating the multimodal capability this thing has to include up to 9 reference of video/audio/image into your prompt to direct your scene.
@maway I understand it can grab things like the persons look, then another image for the background, a third for the mood, the sound of a persons voice, the movements of a video etc.
That doesn't really change my argument. But on the other hand you can argue that it does because you can use the reference basically as a lora.
But I want the chaotic element of the model making a "choice", a reference stays the same.
Agree, what we really need is just propert motion data, female masturbation without a lora right and and theyre just rubbing all over that thing dual wield, and good luck gettin a finger to go in :(
R2V is also massively slower than T2V or I2V
@lost_moon slightly slower I would say. What really kill it is having a video as reference. Now the time it takes is massive.
@zkyt figured that all out on your own, eh?
@lost_moon If u work with images only or audio and image only and resize them to lower size its alright
Ref2va produces noticeably lower video quality than fl2va at the same megapixel and once you start stacking more and more references, you have to lower either the duration or the megapixel to fit them into vram+sysram latent memory.
yes but is slower inference
@Goshtic You can fix this by increasing sampling steps. Ref2va with a lot of image and video refs requires more sampling steps than the 20-25 the default workflow is set to.
This works insane. Pussy pumped vaginas work great :D
tried with img2vid. the pussy just seems to transform on its own. using strenght and clip at 1.00
What did you use to train this? I'm getting pure noise with ai-toolkit.
I used ai-toolkit. I also just go noise but an pretty recent update fixed it for me
@herpderpmerp Thanks, I'll update and try again :)
Greattt work! Please make an penis version! Thanks!
Can i ask what tools and hardware your training on (hoping this can be trained on a 3090 lol)
Rent a Blackwell RTX6000
you can, I have 16Gb vram
@Lady_Valeria I prefer to do my training local lol Great believer in never charging for lora's (no paywalls etc) so renting becomes a no go on my wages.
@herpderpmerp nice! that is a relief after reading the rtx6000 comment lol
AItoolkit? or a different piece of software and are you using videos or images to train? (if videos i assume sliced up and low res? or is that wrong?)
Sorry million questions but the model is too new to have youtube guides and you seem to have a handle on it lol
Thank you for continuing your work on this. Truly a pioneer in the early stages of H3 lora training.
Moar! Lol
Great work. Trying it out soon.
Can't wait to see what we can push with MMH3.
i can't believe that my first lora download from minimax is a pussy
I knew it since it released.
AFAIK you don't need lora for pussy. In r2v just take some pics and it looks good.
you can't be goonin' this hard while also c*ck blocking simultaneously. you're goin' crazy rn.
True but it will make every pussy look the same as the ref image.
A little variation and surprises from the model is more fun than restricting it
YEah, like... I'm sure we'll get some great LoRAs, but righ tnow, images are enough for it to work with and it understands nsfw stuff natively. There are already a lot of examples of this. Blowjobs in particular, both visually and audibly are better than WAN and LTX and it's not even close.
Thanks for your contribution. what were you toolkit settings? and how many videos in your data set? once again, thanks!
pretty much just the default but with layer swapping so it fits my 16Gb Vram.
I only had 11 3 sec videos, that's why the result is pretty overfit. I think bigger datasets don't work before we get the base model or a training adapter because the training breaks with more time. (just my guess I don't actually know)
@herpderpmerp okay that makes sense. I only got up to 500 step on my 5090 and all the samples turned into noise after 500 steps. Im doing something wrong then.
Edit: Nice, done. This is starting to look very promising! But I HIGHLY recommend you rename this to remove mention of "M****x". They are cracking down on this as it violates the license, and they will probably ask Civit to remove NSFW loras that use their name. Just say "H3".
do the asshole to
this
Yep, anus is the thing h3 thinks isn't necessary.
It's a big improvement over what the stock model does lmao
hey we need to start somewhere, looking forward to it.
Works very well in r2v and has a great, positive impact vs the house model. Definitely gonna use it here and there. Thanks!
How are you getting such smooth skin in your example for 0.2? All of my gens, and the ones from shikharisfree532, have stubble. I've tried all sorts of prompting to remove it, but it's always there.
make an innie version