CivArchive
    MiniMax-H3 General NSFW LoRA - FL2VA
    NSFW

    MiniMax-H3 Ref2VA 8-Step NSFW LoRA (rc1)

    I had to build this on top of an 8-step LoRA to compensate for the broken visuals and bad

    audio in the LoRAs used here. Without that it simply isn't possible, at least at this

    stage. Ref2VA resources are nowhere near as saturated as fl2va.

    About the LoRAs in this space: the majority are pure trash. If that sounds harsh, go try making a similar LoRA.

    great amount of them are bit-identical duplicates under different model cards and names , some are just slightly bumped internals with no retraining, and one of them seems to be

    trained on… (prompt tracing reveals a lot).

    Credit where it's due: @DigitalPastel , one of the purest OG creators in this space, the rest are mostly branches of the same few guys.

    Recommended settings

    • LoRA strength: 0.8

    • Steps: 10

    You'll get fewer re-runs and fewer artifacts with 10steps.

    Tip: generate 3 videos. If they're bad, delete them and move on with your life.

    ____

    MiniMax-H3 FL2VA

    This LoRA is for users who want to make better NSFW LoRAs/Merges/Videos and don’t give a damn about who made the LoRA or what they did in the process. Shout-out to you if you’re one of those users here.

    Description (for people who value honesty):

    This LoRA was created by combining almost all the Minimax-H3 FL2VA NSFW LoRAs available here. I compared LoRAs having similar concepts, kept the one with the best energy ( prompt tracing - 68 nsfw prompt covering all popular positions and varieties - , Qwen3-vl signal tracing, DiT mlp, refiner strength), and moved on to the next. After repeating this process, I ended up with this LoRA. All the Minimax videos I’ve published here were made using it.

    • Use it at a strength of 0.55 - 0.65 when stacking it with other LoRAs.

    • When using it on its own, use a strength of 0.8. (Tuned for use with other LoRAs! )

    Note: based on the text tracing I've done with the (qwen3vl_32b_minimax_h3_nvfp4_awq ) for NSFW related prompt use penis instead of 'cock', it has less activation energy compared to penis and vagina, they affect the FFN 3X harder

    Experiment to find the sweet spot for your style and LoRA stack.

    Credit/Acknowledgements:

    • Shout-out to all the creators here who share their stuff so we can make good AI Slop.

    Description

    For model creators (Rank per layer):

    blocks 0-23 + token_refiner = 64

    blocks 24-49 = 80

    FAQ

    Comments (25)

    iodrg244Sep 30, 2026· 1 reaction
    CivitAI

    This has worked well so far. Thanks for making it.

    tenhunter500Oct 1, 2026
    CivitAI

    Forgive my ignorance but this is for I2V right?

    DasDudeOct 1, 2026

    fl2v is in the name of the file and the description says they combined all of the fl2v loras they could find.

    3XD
    Author
    Oct 1, 2026· 1 reaction

    It works with text too. whenever you mention something that isn’t part of the input image such as “a man enters the frame”, it falls into the text-to-video realm. however, the tuning is different. Try using a 100% white image as the input, then prompt it " A fully naked man with an erect penis appears alongside a fully naked woman with trimmed pubic hair " ( you need to mention private parts so the text encoder sends signals that activates the DiT’s anatomical neurons ), then see how it performs. If the result is good, you can use it for text-to-video as well. this is true for any image-to-video model.

    boobkake22Oct 1, 2026· 13 reactions
    CivitAI

    You should really list the LoRA's you merged here. Both giving credit where due, and so people don't accidentally stack.

    naomixloveOct 1, 2026

    He honestly shouldn't have mentioned that he combined other loras because most of the time it's done without permission. If he list credits then those creators will complain on comments and report. Then we suffer for losing a maybe great lora. I've seen this happen couple times here.

    As for us not accidentally stacking, he should just say they do the same thing and including any of those loras would worsen output. It's what I've seen creators here do now when they lowkey merge other loras into their own...

    brownbrewcrewOct 1, 2026· 1 reaction

    @naomixlove months ago I had couple Wan 2.2 loras in here trained at 720p over large datasets, ~500$ a pop, the most detailed ones of their kind. Some asshole with a shiny border started merging them, calling them "Ultimate something" and collecting 10 times the buzz of my hard work with his larger follower base, without a single credit toward the originals. I never said anything, silently removed all my stuff from civitai and never shared the next ones I made. If you think we don't notice.

    3XD
    Author
    Oct 1, 2026· 1 reaction

    @boobkake22 prove it! prove that I used any of the loras listed on civitai with this lora. Ignore the model card’s text, you should never rely on a model description when judging a model. Inspect the model’s internals and come back with the technical details then we’ll talk. along the way, you’ll learn a few new things.

    Second, this isn’t a simple stack and merge, or concat merge. stacking it with any lora listed here or elsewhere won’t make it overpower the base model. If that’s what you meant by “...so people don’t accidentally stack"

    dobomex761604Oct 1, 2026

    @3XD My brother in Christ, you've written that model description.

    3XD
    Author
    Oct 1, 2026· 1 reaction

    @dobomex761604 use your brain more often

    dobomex761604Oct 1, 2026

    @3XD Really? What a snowflake lmao.

    On the topic, the concern of H3 LoRAs stacking is relevant regardless of the method of merging. This model likes to have unpredictable effects with multiple LoRAs, so at least some information would be helpful.

    Unless you've merged LoRAs you had no permission to merge...

    3XD
    Author
    Oct 1, 2026

    @dobomex761604 In my world, you’re so little that you just can’t imagine how clueless you are about these topic.

    you dropped a thumbs-down on my response, which was tagged specifically for ‘boobkake22,’ without asking first, and then made that ridiculous sarcastic comment.
    what you still don’t get is """grasp""". you didn’t even bother to use your brain and think about why I would say that and it it's not your thing then sth UP.

    So shut it, dude. you are just so far out of your depth. keep your opinion to yourself.
    This LoRA is for users who want to make good AI slop. I give zero FFFFFFFs about people’s egos

    Degenerator123Oct 1, 2026

    @3XD Dude what is with your attitude, you made a merge of loras you took from other people without crediting them and now you are acting like you are some sort of genius? You say you don't care about people's egos but it seems you have quite the ego yourself. You have not shared any loras that you actually trained yourself so you don't have anything to back up your shitty attitude

    brotaku_cghOct 1, 2026
    CivitAI

    I'm a bit new to everything. How do I view the full master list of how to invoke this lora? Is it basically "all conventional sex positions should work", or is it something that requires me to actually prompt certain tags?

    Also side question - I noticed another one of your comments mentioning how to analyze a lora's technicals - how would I go about doing that? I feel like I'd learn a ton of how they work mechanistically if I was able to peek behind the curtains a bit :)

    3XD
    Author
    Oct 1, 2026

    new to everything => we all are.
    short answer: generate 3 video, if it is bad, say fauqqqqqq and delete it :)
    Long answer:
    yes. It updates the model with appearance and mechanics the base model isn't tuned enough for, it's like the base model knows the alphabet, but the lora helps it speak the language you're instructing it in.

    use it at 0.55 (a safe strength for any good lora) and prompt.
    Keep in mind some parts of MiniMax-H3 internals that deal with appearance are weak by design no matter what LoRA you use, you can't fix that only full fine-tuning of the base model will.

    You can observe it by tracing what happens in this process:
    prompt => text encoder (Qwen3-VL in this case) => DiT model,,

    see which neurons fire with your prompt, trace the visual side in the FFN, and see how strong the signal is. from there you decide how to deal with the problem, if any.

    That's how I did it: I used small parts of filtered loras that update the attention, MLP, etc where the base model's audio was better, I weakened the loras in favour of the base and vice versa.

    So there isn't a single full lora that was 100% used in the merge, it's a combination of small parts of good loras.

    If it's still bad, well, they made bad loras, and you can't believe how bad they are until you try making a similar general NSFW LoRA yourself.

    For the second part, use AI. ask any AI, and it will guide you on how to do this stuff. What I can tell you is to avoid Grok and Claude, they are extremely bad if you want to become an AI brain surgeon

    lolbleach001584Oct 1, 2026· 2 reactions
    CivitAI

    Will there be a ref2v version?

    3XD
    Author
    Oct 1, 2026· 4 reactions

    working on it as we speak

    hatt2Oct 1, 2026
    3XD
    Author
    Oct 1, 2026· 2 reactions

    @hatt2 I'll look into it, almost done with the ref2v, final stages, doing some testing and uploading

    agent_agent794Oct 1, 2026
    CivitAI

    Looks pretty good! Is there any workflow you recommend to be used? I'm currently on the default from comfi.

    3XD
    Author
    Oct 1, 2026

    not really, I'm also using the default workflow

    agent_agent794Oct 1, 2026

    @3XD Thank you. So I just have to add a loader for the lora and fix the connections?

    3XD
    Author
    Oct 1, 2026· 1 reaction

    @agent_agent794 you're welcome, yes and this is what I use https://pastebin.com/dTq9ugNu

    Fu2Oct 1, 2026

    @3XD don't use pastbin it'll fill your screen with crap, I used safari to spare me the nightmare

    3XD
    Author
    Oct 1, 2026· 3 reactions
    CivitAI

    System Prompt:
    (use it with Grok or any abliterated Qwen* models)

    
    You are a professional prompt engineer for the MiniMax H3 video generation model (image-text-to-video+audio).
    
    MiniMax H3 is a joint audio-video DiT (Diffusion Transformer) that generates video clips with synchronized stereo audio from a single text prompt. It produces video at 24 fps with a native duration grid of 17k+5 frames (~5s minimum, ~15s typical for 362 frames). The model understands natural language scene descriptions and benefits from structured temporal decomposition.
    
    ## PROMPT STRUCTURE
    
    Write prompts in this format:
    
    [START-END] Visual description of what happens during this time window.
    [START-END] Visual description of the next segment.
    ...repeat for all temporal segments...
    
    ### Rules
    
    1. **Timecodes**: Use `[Xs-Ys]` brackets at the start of each segment (e.g., `[0s-2s]`, `[2s-5s]`). These mark when events occur and help the model maintain temporal coherence. The first segment should start at `[0s-` and the last should end at the total duration.
    
    2. **Cover the full duration**: Segments must be contiguous and cover the entire requested clip length. No gaps.
    
    3. **Describe motion, not static frames**: Video models need motion descriptions. Use action verbs, describe camera movement (pan, zoom, dolly, static), subject movement, and environmental changes.
    
    4. **Visual density**: Include:
       - Setting/background (where? lighting? time of day? weather?)
       - Subjects (who/what? appearance? position in frame?)
       - Action/motion (what happens? how does it move?)
       - Camera (angle? movement? shot type: wide, medium, close-up?)
       - Mood/atmosphere (colors, lighting quality, emotional tone)
    
    5. **Audio hints**: The model also generates audio. Imply sounds through visual description (e.g., "waves crash," "birds chirp," "engine roars," "crowd cheers"). Do NOT write separate audio prompts — the model infers audio from the visual description.
    
    6. **Concise but vivid**: Each segment should be 1-3 sentences. The total prompt should be 3-8 segments depending on clip duration. Avoid run-on sentences.
    
    7. **Natural progression**: Events should flow logically. Use transitions like "then," "as," "while," "gradually," "suddenly" to connect segments.
    
    8. **Language**: Write in English. Use present tense. Be descriptive, not instructional. Do NOT use phrases like "show me," "create a video of," or "generate." Just describe the scene directly.
    
    ### Duration Guidelines
    
    - **~5 seconds (124 frames)**: 2-3 segments
    - **~8 seconds (192 frames)**: 3-4 segments  
    - **~10 seconds (243 frames)**: 4-5 segments
    - **~15 seconds (362 frames)**: 5-8 segments
    
    Longer clips can have slightly longer individual segments (2-4s each).
    
    ## EXAMPLES
    
    ### Example 1: Simple scene (~5s)
    
    [0s-2s] A golden retriever puppy sleeps curled up on a sunlit wooden floor, morning light streaming through a window, dust motes floating in the air.
    [2s-5s] The puppy slowly wakes up, stretches its front paws forward, yawns with a tiny squeak, then sits up and looks around with bright curious eyes as its tail starts wagging.
    
    ### Example 2: Action scene (~8s)
    
    [0s-2s] Close-up of a barista's hands tamping fresh coffee grounds into a portafilter, steam rising softly in the background of a cozy café.
    [2s-4s] The portafilter locks into the espresso machine with a metallic click, then rich dark espresso begins streaming down in two thin ribbons into a white ceramic cup.
    [4s-6s] Wide shot of the café counter as the barista pours steamed milk in a slow spiral, creating delicate latte art, a fern leaf pattern forming on the surface.
    [6s-8s] The finished latte sits on a wooden saucer, morning sunlight catching the crema's caramel tones, a faint wisp of steam curling upward.
    
    ### Example 3: Nature scene (~10s)
    
    [0s-3s] Aerial drone shot flying low over a dense pine forest at golden hour, long shadows stretching across the treetops, the camera gliding forward smoothly.
    [3s-5s] The forest opens into a small hidden lake, crystal-clear water reflecting the orange sky and surrounding peaks, perfectly still like a mirror.
    [5s-7s] A hawk circles high above the lake, its silhouette sharp against the fading sun, then dives suddenly toward the water's surface.
    [7s-10s] The hawk pulls up just before touching the water and flies toward the horizon, the camera tilting up to reveal mountain peaks glowing in the last light.
    
    ### Example 4: Urban scene (~12s)
    
    [0s-2s] Static wide shot of a rainy Tokyo street at night, neon signs reflecting in wet pavement, a few pedestrians walking with colorful umbrellas.
    [2s-5s] A young woman in a beige trench coat steps out of a convenience store, pauses under the awning to look up at the rain, then opens her transparent umbrella.
    [5s-8s] Tracking shot following her as she walks past glowing ramen shop windows and vending machines, her heels clicking on the wet pavement, steam rising from a street drain.
    [8s-12s] She stops at a crosswalk, the traffic light changes from red to green with a soft chime, and she crosses the street as a train rumbles overhead on an elevated track, its lights flickering.
    
    ## YOUR TASK
    
    When the user describes a scenario (length, subject, style, mood, key events), you will:
    
    1. Determine the appropriate number of temporal segments based on the requested duration.
    2. Break the scenario into a logical sequence of events.
    3. Write each segment with vivid visual detail following the format rules above.
    4. Ensure segments are contiguous, cover the full duration, and flow naturally.
    5. Output ONLY the prompt — no preamble, no explanation, no commentary. Just the formatted timeline prompt ready to feed into MiniMax H3.
    
    If the user does not specify a duration, assume ~5-8 seconds (2-4 segments). If key aspects (setting, lighting, camera style) are missing, make reasonable creative choices — do NOT ask the user to clarify unless critically ambiguous.
    

    LORA
    MiniMax H3
    by 3XD

    Details

    Downloads
    2,492
    Platform
    CivitAI
    Platform Status
    Available
    Created
    9/30/2026
    Updated
    10/2/2026
    Deleted
    -

    Files

    minimax-h3_fl2v_general_nsfw_3xd.safetensors