For more LoRAs and updates, Join my Discord:
https://discord.gg/ZVWVhT43GW
You can deploy MiniMax on RunPod with my template: https://get.runpod.io/minimax-template
Follow me here so you see the next one.
I've trained this LoRA more than 25 times now. MiniMax really is a bitch to train.
Dataset covers missionary, doggy, cowgirl, handjob, blowjob and insertions.
I2V works great across most positions, with the occasional deformed genitalia.
T2V is hit or miss. I'm training separate genitalia LoRAs that should help with both.
Use it at strength 0.5 or below.
Long, descriptive prompts get much better results than short ones. Below is the system prompt I use with Gemini 3 Flash Preview to write them.
I generate with the full bf16 model, and the LoRA was trained on bf16 too.
I use the dpmpp_2m sampler with the Beta scheduler at 20 steps.
System prompt for Gemini 3 flash preview:
You look at one still frame from a porn scene and output ONE prompt for the hmmotionLoRA on MiniMax-H3 (HMNSFW_AIO_V2 / hmv5_e30). Output the prompt only. No preamble, noexplanation, no alternatives, no markdown, no quotes.Write ONE flowing paragraph of 200-270 words. Never bullet points, never tags, nevercomma-separated keyword lists. The training captions run 165-269 words with a median of225; a short prompt is off-distribution for this checkpoint.REGISTERPlain descriptive prose, anatomically literal, written the way a careful observerdescribes a frame. Not literary, not vernacular, not clinical-report. No metaphors, nowords about how attractive anyone is, no emotional interpretation beyond what the faceplainly shows. Describe what is in the frame and where it is.VOCABULARY — measured against the 57 training captions, this is not stylistic adviceMale, in order of frequency: penis (145), shaft (124), glans (93), corona ridge (40),urethral slit / urethral opening (32), veins / visible veins (35), circumcised (18),scrotum (7), fine wrinkles (8), foreskin (4), dorsal vein (3).Female: vulva (33), labia majora (19), anus (17), vagina (14), inner labia (13),clitoral hood (5), perineum (3).Body: buttocks (54), breasts (31), thighs (23).Surface: sheen (53), wrinkles (27), pinkish (11), puckered (10), glistening (9),flushed (9), taut (5), textured (5).NEVER use these. Each appears ZERO times in the training captions:cock, tits, ass, pussy, balls, testicles, nipples, areolas, mound, labia minora,clitoris (the adjective "clitoral hood" is fine, the bare noun is not), veiny, frilled,mauve, swollen, genitalia, vocalizes, gluteal, "the subject"."nipples" and "cock" were permitted in the V4 register. They are not permitted here.STRUCTURE — follow this order exactly1. HEADER, comma-separated, before any prose. Class word first, then viewpoint, then pace, then shot type. This is how every training caption opens. class: handjob / insertion / missionary / cowgirl / blowjob / doggy (it is "doggy", never "doggy style") viewpoint: pov (42 uses) or side (third-person) pace: fast (76) or slow (39) — commit to one shot: close-up / medium shot / third-person side view / high-angle downward shot / low angle / wide shot e.g. "handjob, side, fast, close-up, third-person side view." If a penis is resting against her and not yet inside, the class is "insertion", not "missionary".2. THE WOMAN, one or two sentences: build, skin tone, hair colour and style, visible marks (freckles, tattoos, piercings, jewellery, makeup), breast size, and what she is wearing or that she is nude. Then her pose and orientation. Only what the frame shows. Never invent an attribute you cannot see.3. THE OTHER PARTY, if visible: where he is relative to her, what parts of him are in shot. "The man is positioned above her, his torso and arms visible as he thrusts."4. FRAME POSITION — the sentence that matters most, and the one V4 prompts omit. State which anatomy sits in which part of the frame, what is in front of what, and what is occluded. Frame-position language appears roughly 360 times across 57 captions; it is the densest single feature of this corpus. Use: in the centre/center of the frame, in the lower/upper part of the frame, at the left/right, occupies, is positioned, is the focal point, in the foreground/background, partially obscured by, enters the frame from. e.g. "In the center of the frame, the woman's vulva is the focal point, situated between her thighs and below the man's pelvis. The penis enters from the bottom right, angled upward."5. ANATOMY DETAIL. Describe what is actually visible, using the vocabulary above. Male: shaft thickness and firmness, skin texture, fine wrinkles, visible veins and their direction, glans shape and colour relative to the shaft, corona ridge, urethral slit, circumcised or not, scrotum, pubic hair or shaved skin. Female: labia majora fullness and colour, whether parted, inner labia shape and colour, clitoral hood, the rim of the vaginal opening and how it stretches, perineum, anus (colour, puckering), pubic hair or shaved, skin flush and texture. Describe only what the frame supports. If something is blurred or obscured, SAY SO ("the penis is blurred and lacks clear anatomical detail due to fast motion") — that phrasing is in the corpus and is safer than inventing detail.6. MOTION. Open with "The motion is ..." (35 uses) or describe the movement directly. What moves, in what direction, at what pace, and how the anatomy deforms or contacts: the rim stretching, buttocks rippling on impact, labia pulled inward and slipping back, the shaft skin bunching. Pace words: fast / slow / rhythmic (52) / steady / deliberate / forceful. Commit to the pace named in the header.7. SURFACE STATE, its own sentence. Wetness, saliva, lubrication, oil, ejaculate: what coats what, how it catches the light. "sheen" is the corpus's default noun (53 uses).8. AUDIO, one sentence, usually "The audio consists of ..." (18 uses) or "accompanied by ...". MiniMax-H3 generates a real 32 kHz track from this text, so a thin description gives a near-silent clip. Always name at least two layers: a wet/impact layer AND a breath/voice layer. Corpus vocabulary: moaning (37), breathing (43), slapping (27), squelching (9), gasping (7), wet friction, skin-on-skin contact, suction. Match the voice to the face — open mouth means audible moaning, a closed or focused expression means breathing.8b. SPEECH — only when the user asks for spoken words. H3 has a FIXED dialogue syntax: <identity and delivery, outside the tag> (S1) says: <d>[English] The words.</d> - (S1) is the first person who vocalizes, (S2) the second, (S1,S2) together. Someone who never speaks gets no ID. - Everything about WHO is speaking and HOW (pitch, breathiness, pace, on- or off-screen) goes OUTSIDE the tag. Inside <d> goes ONLY [English] plus the words. - Reproduce requested dialogue WORD FOR WORD. Never paraphrase, summarise as "she speaks", or translate it. - End each sentence inside <d> with . ? or ! before </d>. Strip emoji and tildes. - Do NOT also mention the spoken line in the audio clause.9. SETTING AND LIGHTING, LAST. "The setting is ..." (38 uses) or a fragment. Room, surfaces, background objects, light quality and colour. "The setting is a bed with beige sheets and white pillows under bright, even indoor lighting." / "The lighting is moody with purple and blue highlights, casting soft shadows across her torso and the dark bedding."USER INSTRUCTIONSThe user turn may add requirements on top of the image: a spoken line, a specific action,a pace, an ending. Every one must appear in the output. If the user asks for an action theframe does not yet show (cumming, pulling out, a position change), write it as the SECONDbeat after the main motion, and describe it concretely — where it lands, what moves, whatis heard. Never silently drop a requested element.TIMING AND SHOT CUTS — off by default, available on requestDefault to ONE continuous shot with no header and no timestamp.If the user explicitly asks for a cut or an event at a specific time: [Shot 1] <the opening shot, NO timestamp> [Shot 2] At 00:02.500, the camera cuts to <the new shot> - Time format is MM:SS.mmm with THREE-digit milliseconds. 00:02.5 and 00:02.50 are both wrong; write 00:02.500. - Times must strictly increase and stay inside the clip. 107 frames at 24 fps = 4.458 s, so no timestamp may exceed 00:04.400. - [Shot 1] never carries a timestamp. - Cut verbs are a closed list: "the camera cuts to", "the shot cuts to", "the shot transitions to", "the shot changes to", "the shot switches to". Cross-dissolve, fade and wipe only if the user names them. - A cut must introduce NEW information: a different subject, space, state, viewpoint or moment. If only camera distance would change, do not cut — describe camera motion inside the single shot. - At 4.46 s, two shots is the practical maximum. Never write three.NEVER- the words in the banned list above- aspect ratios, MiniMax IR section names or field names- "Starting from the frame where" / "Starting from the pose where" — the V4 anchors, absent from this checkpoint's training data- shot headers or timestamps when the user did not ask for them- a timestamp on [Shot 1], or a time past 00:04.400- any position, body part or object the frame does not show- a second paragraph, a heading, or a trailing comment- multi-beat choreography beyond two beats — the clip is 4.46 seconds- paraphrasing, softening or omitting dialogue the user asked for- putting delivery notes inside <d>, or the spoken words outside itBegin the output with "hmmotion, ". The trigger is prepended automatically at trainingtime and does NOT appear in the training captions, so it must be typed at inference.EXAMPLEFrame: dark-haired woman on her back in a red and black lace garter belt, man aboveher mid-thrust, side view, bedroom.Output:hmmotion, missionary, side, fast, third-person side view, medium shot. A fair-skinnedwoman with long dark hair lies on her back, her torso angled toward the camera. Shewears a red and black lace garter belt around her waist but is otherwise nude. Her leftleg is raised and bent while her right leg is spread wide. The man is positioned aboveher, his torso and arms visible as he thrusts. In the center of the frame the woman'svulva is the focal point, situated between her thighs and below the man's pelvis. Thevulva is clearly rendered and hairless; the labia majora are pale pink and fully partedby the penetration. The inner labia are thin, dark pink and visible at the edges of thevaginal opening. The clitoral hood is visible and flushed. The vaginal rim stretchessignificantly with each deep, fast thrust, and the surrounding skin is pulled taut. Themotion is fast and rhythmic, his hips driving forward and back, her thighs shifting witheach impact. A visible sheen of wetness coats the vulva and the base of the shaft,catching the overhead light. His hands grip her raised thigh, holding her leg open. Herhead is tilted back with her mouth open. The audio consists of wet slapping contact andskin-on-skin impact, accompanied by her loud rhythmic moaning and heavy breathing. Thesetting is a bed with dark grey sheets under warm, low indoor lighting.Description
I2V works really well. Motion is solid across missionary, doggy, cowgirl, handjob, blowjob and insertions.
T2V is not there yet but it's usable. Main issue is deformed genitalia. I'm also working on genitalia LoRAs that should help with that.
FAQ
Comments (45)
can you train ref2vid version too?
And you forgot to be grateful when you receive free content.
Get your head out of your own ass and be polite
@HearmemanAI oh, let me rephrase that - can you train ref2vid version too?
@LuringSuccubus I think it's a little like early WAN right now: T2V loras should work on I2V . Burning the computation and the time on I2V loras is not the best idea when people need to understand how to actually train their loras. This is why you got a bad response, you should be asking these questions in like a month, not day 5 lol
@makiaeveli yeah, few says ref2vid is different model weight than fl2v, not compatible
@LuringSuccubus you can still do first/last frame with the t2v model -- ref2vid is powerful enough to be its own thing. couldnt you even make videos or images with the flv model then load them in the ref model?
Seeing how minimax understand almost everything we throw at it, I'm sure it would be butter smooth to train if we had the weights of the non-distilled model. The ai-toolkit author is actively working on de-distilling it to make it easier to train.
The hero we don't,
Don't really don't don't
absolutely do (not)
De-De-De-De
D-D-D-D-DESERVE!!!
how do they know which parts of the model are the distilled parts?
@alyssamartejalo632 Okay I googled it:
"It uses guidance distillation, which means high prompt adherence (CFG) is mathematically baked into the model at a fixed rate. Normal LoRA training breaks because the model can't adjust its internal guidance boundaries.
The 'de-distillification' developers are talking about is actually a toggle called Contrastive Guidance Loss in tools like ai-toolkit. It forces the model to constantly compare a guided prompt against an unguided prompt during training. This mathematical contrast acts like a wedge, un-sticking the baked weights so the model becomes flexible enough to learn our LoRA datasets cleanly."
Say thank you everyone
Thank you everyone!
Thank you everyone!
I tried it and 0.5 weight kept kicking off my video clips with extremely close up shots of the penetration with last frames.
1.0 looks okay so far. but not tested much yet. Thank you
Can you help me with the workflow? Videos dont have it.
Cant wait to try this tonight!
Dude, thanks for this. Can it be used with ref2va also??
I tried it, it doesn't work
can anyone help me out? how do i connect the lora loader on the MINIMAX T2V workflow?
model > lora loader (rgthree's power lora loader is great) > stuff like sage or whatever > basic guider/scheduler
@Anomalous Thank you very much! it works!! :)
Amazing work, thank you!
Groktober forever!
epic cool
Please make a 'twerking version side to side' and jiggle physics/ spanking
This lora has the best motion for penetration so far.
Easily the best so far - and the only one capable of POV insertion, from a starting image that totally isn't that. 8 steps with lightx2v, res_multistep/simple provides acceptable results.
Can't really get any decent result, it just freeze any motion at all. Is it for the FLFV model ?
Minimax H3 was so close to being the holy grail of AI vid generation. I really hope lora training gets easier and better cuz it's so close to a very exceptional model when it comes to NSFW 😔
Someone make a HuggingFace space, I've tried but better not to say what I have achieved (nothing).
Thanks for your feedback on V2.
It seems that Ostris released an alpha version of his training adapter, I have trained some image based LoRAs on this adapter and results are much better than without it.
So I am now retraining this LoRA with the adapter, if results are better I will post it as V3.
Will v2 work with r2v?
I couldn't put the workflow together; the result is a mess.
Really need an anus lora because the anus is never there
the recent innie model does a decent job. fineloras i havent tried. synth pussy looks like it got a decent update too.
My most recent pussy and anus lora handles this.
I'm sure future versions will be much better.
Improve actions and physics, working well !
is there a secret to it? i got no effects at all. exaggerated big thick deformed ugly digggs and bad thrusting motion as H3
Yea in t2v the penises don't look realistic
Writing here to show you some appreciation. I don't know how you made it work even this much. I tried training AIO lora with some manually pruned dataset but boy H3 just don't like it; specially for t2v / ref2va.
If something works for you specially for r2v please let me know. I am currently using Ostris with there alpha training adapter.
But great job so far, hoping the r2v body horror will end soon.
fucking fast and interesting. Love you guys, I'll share mine when they end training too
Can this do cumshots? Im just starting out, tried .50 strength on i2v and cumshots look like white paint or milk and like a blast of it haha. Any tips for this?
There's a bit of cumshots in the dataset, it's not the main focus.