CivArchive
    Tiktok Catgirl Cosplay - MiniMax H3 (FL2VA) - v1.0
    NSFW

    Tiktok Catgirl Cosplay - MiniMax H3 (FL2VA)

    Trained on real cosplay footage, so it pushes H3 towards looking like an actual photo of someone in a wig and ears instead of the smooth idol face it falls back on. Skin texture comes back, lighting gets harder, makeup looks applied rather than rendered.

    H3 already knows what a catgirl is so this isn't teaching it that. It mostly changes the look. It does follow colour details better though, ask for crimson inner ear fur and you get crimson instead of pink.

    Trigger is `ctgrlcos`. Put it mid sentence where you'd normally name the person, not stuck on the front.

    Settings

    Strength 1.0. That's the only value I've actually tested, I only ever ran it on and off to check it was working, so if 0.7 looks better to you then use that.

    - FL2VA, pruned checkpoint. Samples were made on pruned bf16, int8-convrot and nvfp4 should be fine too.
    - CFG stays at 1. The weights are distilled, there's no negative prompt.
    - res_multistep + simple, sigma shift left alone at 12.0 / 3.0.
    - 768x1344 portrait is what it wants, that's the biggest bucket it saw.
    - Frame count has to sit on the 17k+5 grid, so 124 for about 5s or 243 for 10s.

    Prompting

    Captions were written in H3's three field format so it responds best to the same shape:

    integrated_multimodal_description: [Shot 1] Live-action, Medium close-up, ctgrlcos wears
    large fluffy white cat ears with black tips on a thin black headband, her long straight
    platinum-blonde wig falling past her shoulders. She wears a white collared shirt with a red
    bow. She tilts her head and smiles in warm lamplight. The camera holds a static shot.
    
    overall_soundscape: N/A
    
    non_diegetic_music: N/A

    They ran long, 90 to 160 words, always opening with [Shot 1] Live-action and the shot type and ending on a camera line. Short prompts still work, you just get less out of it. Audio fields were N/A on all of them since it's a stills dataset, so it won't help you with sound.

    Careful with unpruned checkpoints

    Load this on an unpruned model and 50 of its 258 modules quietly don't apply. Pruning changes the AdaLN input width to 8 and this was trained on a pruned base, so those shapes don't line up. Nothing visibly breaks, you just get a weaker result and no obvious reason why.

    Easiest check is the ComfyUI log when the model stages, you want `258 patches attached`. If it says 0 then the lora isn't doing anything at all.

    Ref2VA loads, the key shapes are identical, but I haven't properly looked at whether the output holds up. Assume untested.

    Known issues

    • Trained on still images, used for video, so fine detail can come out soft. Main thing I may fix if I make a next one.

    • Most of the dataset is medium close up. Strong on face shots, falls off once the subject is small in frame.

    • Backgrounds drift towards bedrooms and plain walls even when you prompt otherwise. One of the sample clips does exactly this, I asked for a plain studio backdrop and got a room.

    Samples and dataset

    Each clip is the same seed and prompt run twice, lora off on the left, 1.0 on the right. Worth knowing the trigger is in both prompts, so the left side is just H3 dealing with a word it doesn't know.

    Dataset is 248 still short form cosplay video, filtered down to cat ears only since a lot of the source was fox and bunny, captioned locally with Qwen3-VL-32B, then I went through them by hand and cut it from 289 down to 248.

    Rank 16, 3000 steps, ai-toolkit, about 48 minutes on a 5090 (without sample output).

    Description

    FAQ

    LORA
    MiniMax H3

    Details

    Downloads
    208
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/10/2026
    Updated
    8/10/2026
    Deleted
    -
    Trigger Words:
    ctgrlcos

    Files

    h3_catgirl_cosplay_v1.safetensors

    Mirrors