LTX2.3/MiniMax Free Generation!
I turned off generation charging. I didn't know it was on must have been a autojoin thing. I may start doing early release just to help cover some small costs but not per generation.
Deepthroat
What's new
Rank 16.... 1/4 the size of the original
Dataset prompt rewrite, a full rewrite of prompts for Visual and Audio
Audio. Its still fucking AI audio its hit and miss is vastly better now.
Images. Better quality dataset.
Eye contact.
maintain eye contactnow works as a prompt phrase and holds through the shot mostly depends on the angle.
Better side profile. More training data around it.
Length. Drives a full 15 seconds, where earlier versions lost the action around 8.
Multiple people. MMF and FFM interactions are supported, you will need to figure out the prompting.
Entering the Frame. Better supported.
More thrusting, with more variation in pace and some actual in and out with needing extra loras.
Better prompt adherence overall — and it's the reason the usable strength window is as wide as it is.
Sample generations
Everything shown in gallery was generated at 0.4 MP with no post-processing — no upscaling, no interpolation, no audio cleanup. Workflows are included with their custom prompts, so results are reproducible as-is.
Visual vocabulary
Measured across the 57 training captions. Coverage is the share containing the term — things bind reliably from about 15% upward and fall apart below 10%.
Anatomy — the word choices that matter
| Use | Coverage | Not |
|---|---:|---|
| shaft | 77% | — the single best-covered anatomical term |
| penis | 75% | dick 2%, cock 5%, member 0% |
| glans | 49% | tip 2%, head of 0% |
| scrotum | 0% | balls 7% — neither is trained. Avoid the region or expect the base model to improvise. |
Fluid and texture
| Term | Coverage | Notes |
|---|---:|---|
| saliva | 51% | |
| glistening | 51% | |
| spit | 16% | Works, but weaker. thick strands of spit |
| wet | 26% | |
| smooth | 26% | |
### Motion
| Term | Coverage | Notes |
|---|---:|---|
| slides | 26% | The core verb. |
| pulls | 28% | |
| reappears / disappears | 32% / 26% | The in-and-out cycle. Naming both gives you rhythm control. |
| bob | 16% | |
| thrust | 14% | |
### Face and gaze
| Term | Coverage | Notes |
|---|---:|---|
| lips | 74% | |
| mouth | 61% | |
| eyes | 54% | |
| face | 53% | |
| she maintains eye contact | 18% | Confirmed to hold through the shot. |
| wince | 0% | Not in the captions at all, but the base model looks good |
### Act terms are nearly always on
deepthroat 63%, blowjob 28%. Both are close to constants across the dataset, so they anchor the model in the right place but give you little steering. Adding them costs nothing; expect them not to change much.
Shot, light and setting
| Term | Coverage | Notes |
|---|---:|---|
| hair | 90% | Near always-on. |
| static / steady camera | 19% / 18% | handheld 11%. |
Audio Vocabulary
These five were fixed across the whole training set and carry enough examples to be steerable. Use the exact word — synonyms were deliberately trained out.
sucking · squelches · gagging · moaning · breathing
| Term | Coverage | Meaning as trained |
|---|---:|---|
| sucking | 72% | Soft continuous wet mouth sound. The bed the rest sits on. |
| moaning | 86% | Voiced non-verbal vocalisation. |
| breathing | 72% | Audible breath, panting, sharp inhale or exhale. |
| squelches | 41% | Sharp, airy, individually countable suction-release events. |
| gagging | 30% | Throat rejection sound. Weakest of the five — repeat it or pair it with intensity words. |
Words that do nothing in the lora (maybe the model does something with them)
~gawk gawk~~ ~slurping~~ wet mouth sounds
gawk gawk is worse than useless — the text encoder reads it as staring. The last two were actively suppressed during captioning; they will be ignored or mapped onto sucking.
Steering the mix
Rhythm and intensity are ordinary English and vary freely — how fast, how wet, what sits on top of what, whether the room is quiet. The five terms are the only fixed tokens. Describing the room works and is worth doing; a few clips carry noticeable reverb and naming it gives you control over it.
---
What to expect
Audio level runs quiet. Training material was deliberately left un-normalised, with a median around −26 LUFS, because loudness normalisation raises every noise floor by exactly as much as it raises the content. Output is clean but quiet — normalise after generation rather than fighting it in the prompt.
Sound is genuinely prompt-driven. Rewording overall_soundscape changes the output. The vocabulary bound rather than memorised, so the soundscape field is a real control surface and not decoration.
Seed matters more than you expect. Across fixed prompts, variation between seeds is larger than variation between neighboring checkpoints. If a take is close but wrong, reroll before you rewrite.
---
Nerd shit
A joint audio-video LoRA for MiniMax-H3. Video and sound are generated together from one prompt, not dubbed afterward.
| | |
|---|---|
| Checkpoint | epoch 31 |
| Base | MiniMax-H3 minimax_h3_fl2va_bf16) |
| Mode | fl2va |
| Rank / alpha | 16 / 16 |
| File | 284 MB |
---
Training configuration
| Parameter | Value | |
|---|---|---|
| Base model | minimax_h3_fl2va_bf16 | |
| Network | lora_minimax_h3 | 200 modules across 50 DiT blocks |
| Rank / alpha | 16 / 16 | ~9 directions carry 90% of the energy |
| Training mode | fl2va | first + last frame conditioning |
| Optimizer | AdamW | constant LR 1e-4, no warmup |
| Precision | bf16 | fp8 base, block quantisation |
| Audio loss weight | 2.0 | video 1.0, balanced per modality |
| Guidance distillation | 4.0 | normalized form, sigma schedule |
| Base preservation | 0.02 | evaluated every batch |
| Flow shift | 12.0 / 3.0 | video / audio, off one shared coordinate |
| Timestep sampling | uniform | |
| Seed | 42 | |
### Dataset
Videos and image all 24 fps, across eight bucket resolutions from 288×512 to 1024×576. Clips were culled on measured quality rather than by eye: content-to-noise-floor ratio for audio, and bits-per-pixel plus a resolved-detail test for video. About a third of the video carries no audio track and trains motion only.
Recommended inference guidance scale: 4. Trained at 24 fps — other frame rates are out of distribution.
Description
# v0.2
What's new
Rank 16.... 1/4 the size of the original
Dataset prompt rewrite, a full rewrite of prompts for Visual and Audio
Audio. Its still fucking AI audio its hit and miss are vastly better now.
Images. Better quality dataset.
Eye contact.
maintain eye contactnow works as a prompt phrase and holds through the shot mostly depends on the angle.
Better side profile. More training data around it.
Length. Drives a full 15 seconds, where earlier versions lost the action around 8.
Multiple people. MMF and FFM interactions are supported, you will need to figure out the prompting.
Entering the Frame. Better supported.
More thrusting, with more variation in pace and some actual in and out with needing extra loras.
Better prompt adherence overall — and it's the reason the usable strength window is as wide as it is.
Sample generations
Everything shown in gallery was generated at 0.4 MP with no post-processing — no upscaling, no interpolation, no audio cleanup. Workflows are included with their custom prompts, so results are reproducible as-is.
Visual vocabulary
Measured across the 57 training captions. Coverage is the share containing the term — things bind reliably from about 15% upward and fall apart below 10%.
Anatomy — the word choices that matter
| Use | Coverage | Not |
|---|---:|---|
| shaft | 77% | — the single best-covered anatomical term |
| penis | 75% | dick 2%, cock 5%, member 0% |
| glans | 49% | tip 2%, head of 0% |
| scrotum | 0% | balls 7% — neither is trained. Avoid the region or expect the base model to improvise. |
Fluid and texture
| Term | Coverage | Notes |
|---|---:|---|
| saliva | 51% | |
| glistening | 51% | |
| spit | 16% | Works, but weaker. thick strands of spit |
| wet | 26% | |
| smooth | 26% | |
### Motion
| Term | Coverage | Notes |
|---|---:|---|
| slides | 26% | The core verb. |
| pulls | 28% | |
| reappears / disappears | 32% / 26% | The in-and-out cycle. Naming both gives you rhythm control. |
| bob | 16% | |
| thrust | 14% | |
### Face and gaze
| Term | Coverage | Notes |
|---|---:|---|
| lips | 74% | |
| mouth | 61% | |
| eyes | 54% | |
| face | 53% | |
| she maintains eye contact | 18% | Confirmed to hold through the shot. |
| wince | 0% | Not in the captions at all, but the base model looks good |
### Act terms are nearly always on
deepthroat 63%, blowjob 28%. Both are close to constants across the dataset, so they anchor the model in the right place but give you little steering. Adding them costs nothing; expect them not to change much.
Shot, light and setting
| Term | Coverage | Notes |
|---|---:|---|
| hair | 90% | Near always-on. |
| static / steady camera | 19% / 18% | handheld 11%. |
Audio Vocabulary
These five were fixed across the whole training set and carry enough examples to be steerable. Use the exact word — synonyms were deliberately trained out.
sucking · squelches · gagging · moaning · breathing
| Term | Coverage | Meaning as trained |
|---|---:|---|
| sucking | 72% | Soft continuous wet mouth sound. The bed the rest sits on. |
| moaning | 86% | Voiced non-verbal vocalisation. |
| breathing | 72% | Audible breath, panting, sharp inhale or exhale. |
| squelches | 41% | Sharp, airy, individually countable suction-release events. |
| gagging | 30% | Throat rejection sound. Weakest of the five — repeat it or pair it with intensity words. |
Words that do nothing in the lora (maybe the model does something with them)
~gawk gawk~~ ~slurping~~ wet mouth sounds
gawk gawk is worse than useless — the text encoder reads it as staring. The last two were actively suppressed during captioning; they will be ignored or mapped onto sucking.
Steering the mix
Rhythm and intensity are ordinary English and vary freely — how fast, how wet, what sits on top of what, whether the room is quiet. The five terms are the only fixed tokens. Describing the room works and is worth doing; a few clips carry noticeable reverb and naming it gives you control over it.
---
What to expect
Audio level runs quiet. Training material was deliberately left un-normalised, with a median around −26 LUFS, because loudness normalisation raises every noise floor by exactly as much as it raises the content. Output is clean but quiet — normalise after generation rather than fighting it in the prompt.
Sound is genuinely prompt-driven. Rewording overall_soundscape changes the output. The vocabulary bound rather than memorised, so the soundscape field is a real control surface and not decoration.
Seed matters more than you expect. Across fixed prompts, variation between seeds is larger than variation between neighboring checkpoints. If a take is close but wrong, reroll before you rewrite.
---
Nerd shit
A joint audio-video LoRA for MiniMax-H3. Video and sound are generated together from one prompt, not dubbed afterward.
| | |
|---|---|
| Checkpoint | epoch 31 |
| Base | MiniMax-H3 minimax_h3_fl2va_bf16) |
| Mode | fl2va |
| Rank / alpha | 16 / 16 |
| File | 284 MB |
---
Training configuration
| Parameter | Value | |
|---|---|---|
| Base model | minimax_h3_fl2va_bf16 | |
| Network | lora_minimax_h3 | 200 modules across 50 DiT blocks |
| Rank / alpha | 16 / 16 | ~9 directions carry 90% of the energy |
| Training mode | fl2va | first + last frame conditioning |
| Optimizer | AdamW | constant LR 1e-4, no warmup |
| Precision | bf16 | fp8 base, block quantisation |
| Audio loss weight | 2.0 | video 1.0, balanced per modality |
| Guidance distillation | 4.0 | normalized form, sigma schedule |
| Base preservation | 0.02 | evaluated every batch |
| Flow shift | 12.0 / 3.0 | video / audio, off one shared coordinate |
| Timestep sampling | uniform | |
| Seed | 42 | |
### Dataset
Videos and image all 24 fps, across eight bucket resolutions from 288×512 to 1024×576. Clips were culled on measured quality rather than by eye: content-to-noise-floor ratio for audio, and bits-per-pixel plus a resolved-detail test for video. About a third of the video carries no audio track and trains motion only.
Recommended inference guidance scale: 4. Trained at 24 fps — other frame rates are out of distribution.