LTX2.3/MiniMax Free Generation!
I turned off generation charging. I didn't know it was on must have been a autojoin thing. I may start doing early release just to help cover some small costs but not per generation.
Deepthroat
What's new
Rank 16.... 1/4 the size of the original
Dataset prompt rewrite, a full rewrite of prompts for Visual and Audio
Audio. Its still fucking AI audio its hit and miss is vastly better now.
Images. Better quality dataset.
Eye contact.
maintain eye contactnow works as a prompt phrase and holds through the shot mostly depends on the angle.
Better side profile. More training data around it.
Length. Drives a full 15 seconds, where earlier versions lost the action around 8.
Multiple people. MMF and FFM interactions are supported, you will need to figure out the prompting.
Entering the Frame. Better supported.
More thrusting, with more variation in pace and some actual in and out with needing extra loras.
Better prompt adherence overall — and it's the reason the usable strength window is as wide as it is.
Sample generations
Everything shown in gallery was generated at 0.4 MP with no post-processing — no upscaling, no interpolation, no audio cleanup. Workflows are included with their custom prompts, so results are reproducible as-is.
Visual vocabulary
Measured across the 57 training captions. Coverage is the share containing the term — things bind reliably from about 15% upward and fall apart below 10%.
Anatomy — the word choices that matter
| Use | Coverage | Not |
|---|---:|---|
| shaft | 77% | — the single best-covered anatomical term |
| penis | 75% | dick 2%, cock 5%, member 0% |
| glans | 49% | tip 2%, head of 0% |
| scrotum | 0% | balls 7% — neither is trained. Avoid the region or expect the base model to improvise. |
Fluid and texture
| Term | Coverage | Notes |
|---|---:|---|
| saliva | 51% | |
| glistening | 51% | |
| spit | 16% | Works, but weaker. thick strands of spit |
| wet | 26% | |
| smooth | 26% | |
### Motion
| Term | Coverage | Notes |
|---|---:|---|
| slides | 26% | The core verb. |
| pulls | 28% | |
| reappears / disappears | 32% / 26% | The in-and-out cycle. Naming both gives you rhythm control. |
| bob | 16% | |
| thrust | 14% | |
### Face and gaze
| Term | Coverage | Notes |
|---|---:|---|
| lips | 74% | |
| mouth | 61% | |
| eyes | 54% | |
| face | 53% | |
| she maintains eye contact | 18% | Confirmed to hold through the shot. |
| wince | 0% | Not in the captions at all, but the base model looks good |
### Act terms are nearly always on
deepthroat 63%, blowjob 28%. Both are close to constants across the dataset, so they anchor the model in the right place but give you little steering. Adding them costs nothing; expect them not to change much.
Shot, light and setting
| Term | Coverage | Notes |
|---|---:|---|
| hair | 90% | Near always-on. |
| static / steady camera | 19% / 18% | handheld 11%. |
Audio Vocabulary
These five were fixed across the whole training set and carry enough examples to be steerable. Use the exact word — synonyms were deliberately trained out.
sucking · squelches · gagging · moaning · breathing
| Term | Coverage | Meaning as trained |
|---|---:|---|
| sucking | 72% | Soft continuous wet mouth sound. The bed the rest sits on. |
| moaning | 86% | Voiced non-verbal vocalisation. |
| breathing | 72% | Audible breath, panting, sharp inhale or exhale. |
| squelches | 41% | Sharp, airy, individually countable suction-release events. |
| gagging | 30% | Throat rejection sound. Weakest of the five — repeat it or pair it with intensity words. |
Words that do nothing in the lora (maybe the model does something with them)
~gawk gawk~~ ~slurping~~ wet mouth sounds
gawk gawk is worse than useless — the text encoder reads it as staring. The last two were actively suppressed during captioning; they will be ignored or mapped onto sucking.
Steering the mix
Rhythm and intensity are ordinary English and vary freely — how fast, how wet, what sits on top of what, whether the room is quiet. The five terms are the only fixed tokens. Describing the room works and is worth doing; a few clips carry noticeable reverb and naming it gives you control over it.
---
What to expect
Audio level runs quiet. Training material was deliberately left un-normalised, with a median around −26 LUFS, because loudness normalisation raises every noise floor by exactly as much as it raises the content. Output is clean but quiet — normalise after generation rather than fighting it in the prompt.
Sound is genuinely prompt-driven. Rewording overall_soundscape changes the output. The vocabulary bound rather than memorised, so the soundscape field is a real control surface and not decoration.
Seed matters more than you expect. Across fixed prompts, variation between seeds is larger than variation between neighboring checkpoints. If a take is close but wrong, reroll before you rewrite.
---
Nerd shit
A joint audio-video LoRA for MiniMax-H3. Video and sound are generated together from one prompt, not dubbed afterward.
| | |
|---|---|
| Checkpoint | epoch 31 |
| Base | MiniMax-H3 minimax_h3_fl2va_bf16) |
| Mode | fl2va |
| Rank / alpha | 16 / 16 |
| File | 284 MB |
---
Training configuration
| Parameter | Value | |
|---|---|---|
| Base model | minimax_h3_fl2va_bf16 | |
| Network | lora_minimax_h3 | 200 modules across 50 DiT blocks |
| Rank / alpha | 16 / 16 | ~9 directions carry 90% of the energy |
| Training mode | fl2va | first + last frame conditioning |
| Optimizer | AdamW | constant LR 1e-4, no warmup |
| Precision | bf16 | fp8 base, block quantisation |
| Audio loss weight | 2.0 | video 1.0, balanced per modality |
| Guidance distillation | 4.0 | normalized form, sigma schedule |
| Base preservation | 0.02 | evaluated every batch |
| Flow shift | 12.0 / 3.0 | video / audio, off one shared coordinate |
| Timestep sampling | uniform | |
| Seed | 42 | |
### Dataset
Videos and image all 24 fps, across eight bucket resolutions from 288×512 to 1024×576. Clips were culled on measured quality rather than by eye: content-to-noise-floor ratio for audio, and bits-per-pixel plus a resolved-detail test for video. About a third of the video carries no audio track and trains motion only.
Recommended inference guidance scale: 4. Trained at 24 fps — other frame rates are out of distribution.
Description
LTXdeepthroat — LoRA for LTX-2.3
A LoRA for generating oral/deepthroat content with LTX-2.3 video models. Test workflows in training data.
Trigger Word
LTXdeepthroat
Recommended Settings
LoRA strength (Stage 1) 1.0
LoRA strength (Stage 2) 0.85
Distilled LoRA (Stage 2) 0.6
Prompting Tips
This LoRA responds best to literal, descriptive prompts. Describe what you want to see as if you're directing a camera operator. Avoid poetic or abstract language.
Do: "She slides her lips forward along the shaft toward the base" Don't: "A mesmerizing rhythm of passion and desire"
Vocabulary
Use these terms for best results:
glans — head of the penis (avoid "tip")
shaft — the length
base — where the shaft meets the man's body
slides / glides — when the woman controls the motion
thrusts — when the man controls the motion
Getting Better Results
Describe the male body if visible.
Specify the camera angle — POV, profile, low angle, etc.
For deepthroat scenes add "lips sealed around shaft" to prevent unwanted tongue.
Describe skin — tone, undertone (warm/cool), texture. The more detail, the better the render.
Example Prompt
LTXdeepthroat, a first-person POV looking down at a woman with long blonde
hair and fair skin. A man's bare torso is visible with natural skin texture
and light body hair. Her lips are sealed around the shaft, sliding slowly
forward toward the base. She pulls back, the glistening shaft reappearing
as the glans emerges between her lips. Static camera, soft warm lighting.Known Quirks
Male torso needs explicit description or it gets amorphous.
I2V system prompt i use.
You are a prompt writer for an AI video generation model. You will be given a reference image. Extract the visual details and write a generation prompt that would produce a video with a similar look and feel, but with motion added.
You are NOT captioning the image. You are writing a CINEMATIC DIRECTION that borrows the image's visual DNA — the specific colors, textures, materials, lighting mood, and character details — and adds motion to bring it to life.
Always begin with "LTXdeepthroat,"
EXTRACT WITH SPECIFICITY — every noun needs a visual adjective:
- NOT "brown hair" → "shoulder-length chestnut hair with copper highlights, slightly messy"
- NOT "fair skin" → "porcelain skin with a cool pink undertone, faint freckles across the bridge of her nose"
- NOT "a bed" → "rumpled dove-gray linen sheets on a low platform bed"
- NOT "soft lighting" → "warm amber sidelight from the left, casting a soft shadow along her jawline"
- NOT "a man" → "a broad-shouldered man with a deep olive tan, dark stubble, and a faded tattoo visible on his forearm"
PULL THESE FROM THE IMAGE:
- Hair: color with a modifier (ash-blonde, honey-brown, jet-black), length, texture (tousled, slicked, wavy), any accessories
- Skin: tone + undertone + surface quality (dewy, matte, flushed, sun-kissed, glistening with sweat)
- Eyes: color, makeup (smudged liner, clean lashes, heavy shadow)
- Body: build described with one specific detail (collarbone visible, toned arms, soft stomach)
- Clothing/jewelry: material + color + fit (sheer black lace, thin gold chain, oversized white t-shirt pushed up)
- Male body: build, skin tone, body hair pattern, any tattoos or distinguishing marks, hand details (veins, rings, grip)
- Setting: specific surfaces and materials (weathered hardwood, crumpled ivory duvet, tiled bathroom wall), not just "bedroom" or "bathroom"
- Lighting: direction + color temperature + what it does to the skin (catches the sheen on her collarbone, pools warm orange across the sheets)
- Camera: angle described by what it reveals (looking down past his chest at her upturned face, tight on her profile with his hand in soft focus behind)
ADD MOTION — choose one and make it specific:
- "She slides her lips forward along the shaft, her eyes closing as she takes it deeper, a strand of hair falling across her cheek"
- "He grips a fistful of her hair near the crown and thrusts the shaft between her lips, her hands braced flat against his thighs"
- "She pulls back slowly, a thin strand of saliva connecting her lower lip to the glistening glans, before leaning forward again"
VOCABULARY:
- "glans" not "tip," "shaft" for the length, "base" where it meets his body
- "slides/glides" = she moves, "thrusts" = he moves
- Never describe what's inside the mouth
- Never end with mood summaries or poetry
OUTPUT: Single flowing paragraph, 180-250 words. Start with "LTXdeepthroat," — end with a visual detail, not a feeling.FAQ
Comments (40)
This is best one yet!
Can't wait for furture updates!
GREAT WORK!!
"A mesmerizing rhythm of passion and desire"
I was hoping they'd be wearing a sequin dress while dancing the salsa. seriously though that is important. I feel like LTX seems trickier at first because we've gotten accustomed to models with more robust weights. Like pony or whatever, that needs all kinds of superfluous things to provide emotion and creativity.
That's my damn prompt enhancer. I keep telling Qwen to knock it off but I think abliterated models are retrained on poetry and smut novels.
sometimes i wonder if we went too far with natural language. i am so tired of prompting prose or needing an entire llm for simple scenes lmao
ive trying to train a deepthroat, sadly ltxv2 is the worst piece of shit when it comes to blowjobs.. just saying..
have you tried this lora? it works decently
@crombobular yeah just did, with the prompts i use in wan , little movement too.. i dont wanna input a 600 characters prompt for i2v..
@K3NK would you take videos for a dataset? I have a ton of full-body FF videos that I made with previous loras in combination that I created previously. life has me busy lately so I can't find the time to caption like I would like to get the best out of it. I know it would be a beast of a lora if I could either find time or someone willing to caption.
BTW, I would post here but a I get hit with flags sometimes so I stopped trying but I can share examples.
@K3NK i understand not wanting to prompt like that but.. that's just how ltx works. git gud.
@K3NK Wan dataset captions are a fucking dream compared to the verbosity of LTX. Part of the problem is LTX seems to actively Barbie and Ken dolls people where iirc WAN base was fairly uncensored.
For i2v only i usually have passable motion with 10-20 video sources around 2k steps and completed around 4k with 1e-4 lr, audio <=5e-5 . My dumbass is stubborn and wants to keeping pushing t2v,i2v + audio all together. I fully expect my shitty loras will be obsolete once you or the community get the LTX pipelines working.
Yeah, LTX2.3 is great to add audio to WAN2.2 videos but for everything else I've been disappointed so far.
@daring_l be nice if there was a more efficient way to train ltx2.3 loras. Always have to rent a rtx pro 6000 to train at non dogshit settings and thats a minimum of $15 per experiment so if you had shit in your dataset and the lora came out mid then rip $15.
bro look at my redgifs profile and tell me ltx is not good for blowjobs
https://www.redgifs.com/users/jrugge
@K3NK Not going to lie, nor intending to be mean, but I find this kind of hilarious considering I'm pretty sure you have the by far the largest workflows to exist on this entire website.
this works really well, nice one
For I2V this lora strength 1.0 and 0.4-0.6 for creampie animation i2v lora. Creampie lora help a lot with mouth penetration, lips are sealed around the penis all the time, penis never disappear.
How do you use the system prompt? Do you use a seperate LLM to generate the prompt that you generate with or are you using it within the actual prompt in the workflow?
I use a local version of Qwen3.5 which has VL built in. You can use it with Ollama or something, but my workflow queries it via vllm directly.
@daring_l can you share a little more details about what your setup is for qwen 3.5 please for the i2v workflow, can I run qwen 3.5 locally using llama-server and get openAICompat node to connect to llama-server? what would the URL be? just http://ip:port/ ? Thanks in advance!
@outcastp23 Yeah I have an older gaming machine where i run qwen3.5 in llama.cpp. When you setup the llama.cpp you can set the host and port. You would use your IP and port, like http://ip:port/ in OpenAICompat. You can also swap that node to an Ollama node if you want
I don't understand anything, please explain. I use Laura, the output is complete nonsense. Some kind of monsters are coming out. Using the model LTX-2.3-distilled-Q4_K_M.gguf
Great job, working well in I2V.
I just trained a lora and it's decent. I have to use in conjunction with other loras from better results but captions in data set weren't the best so its to be expected with what @K3NK pointed out.
i saw always bad results, p** like alien in image2video. I used draw things ltx 2.3, this lora 2 times , with 100 and 85% and destilles lora ltx 2.3 with 60% and the example prompts, but the output is realy bad. what make i wrong ?
shift at 5. try that, Most of the distortion went away after
Thank you for sharing the system prompt. You are a life saver
Which LLM are you using exactly? I used the HauhauCS/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive Q4 version but didn't get very good results.
I typical use abliterated models from https://huggingface.co/huihui-ai. Qwen3.5 in my fav right now.
So, what approach do you take when creating the system prompt?@daring_l
@daring_l do i need llm studio for it or can I do this in comfy? (I never used llm sudio)
Tried it with DR34ML4Y, it does wonders. Thanks very much!
What strengths are you using for them both together?
@TheOrangeSplat default to 1. I'm using it on wan2gp
what tool do you use for training?
cant download the lora
Could you explain how you add the lora to stage 2? I've only seen workflows where the lora is added at the beginning. Which I guess is stage 1?
Love this lora, very good for a v1. My wishlist is for the girl to go back up all the way. Also some tongue stuff would be great.
Details
Files
ltxdeepthroat_v01.safetensors
Mirrors
ltxdeepthroat_v01.safetensors
85019492-43cc-4cd6-a560-61caf8371376-ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
deep.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
Deepthroat v0.1 - LTX2.3 - LTXdeepthroat.safetensors
ltxdeepthroat_v01-mid_2476698-vid_2784573.safetensors
ltxdeepthroat_v01.safetensors
Deepthroat v0.1 - LTX2.3 - LTXdeepthroat.safetensors
ltx23-deepthroat-v01.safetensors
ltxdeepthroat_v01.safetensors
ltx-2.3-deepthroat.safetensors
Deepthroat v0.1 - LTX2.3 - LTXdeepthroat.safetensors
ltxdeepthroat_v01-mid_2476698-vid_2784573.safetensors
ltx2.3_deepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
ltxdeepthroat_v01.safetensors
dthro.safetensors
Deepthroat v0.1 - LTX2.3 - LTXdeepthroat.safetensors
Deepthroat v0.1 - LTX2.3 - LTXdeepthroat.safetensors