First test with LTXV23.
Trying to see how LTXV23 handles feet... Not too well as it turns out 😆
The lora does work okay, and definitely improves feets knowledge.
Weight ~1.4
This high strength seems weird to me but per LTX own docs, apparently this is typical
Trigger is "footjob."
See preview videos for other helpful phrases.
Focused on I2V. Cock knowledge is very limited.
Works okay when feet are already in position, but otherwise... good luck.
Please don't ask for tech support. I have no idea what I'm even doing. I'm using the 2-stage distilled workflow from LTX, that's all I know.
Do feel free to leave comments about how shitty this is so I can try to target improvements.
technobullshit:
Trained to 1000 steps. Beyond that things got weird fast, with hands turning into feet and all kind of horror shit.
The training clips were:
five close-ups (just feet and cock) 512x512, 24fps, 65 frames.
four full-body 448x576, 24fps, 89 frames.
I found that 'context' training data (the full-body clips) were extremely important. This is different than I'm used to with Wan which seems to tolerate less context.
I plan to continue working on this, but this seems workable enough for a first version.
Description
FAQ
Comments (8)
Great. Thank you. Can you make just normal feet Lora? Ltx needs to learn bare feet!!
So true. LTX is very bad at feets and soles.
True indeed. I almost gave up before I began when I saw how bad LTX feet were.
Not sure I'll do a "general feets knowledge" lora... trying to train a model what feet look like in any position/action may be outside my ability. But with how well it learned on this one, I might give it a try.
@icelouse Ture hero.
@AndyZocker Use 10eros, like it's a good base anyway where you can throw on concept loras, although not great at raw video generations, but honestly I find I2I to be better anyway because you can much more quickly iterate the starting frame and you can choose whatever image model you want which is epically useful if you're not after realism, although if your not after realism and you are going to significantly change the prospective Dasuwa is probably better but the sound sucks as far as I have tested the model, I don't mean talking, but like lewd noises, mind you 10eros isn't amazing or anything, but it's passable if you can't be bothered to source you're own audio for conditioning.
Wow, the lora is so small! Do you think a larger lora could improve the foot understanding?
This appears to be the typical size for a rank 16, video-only, bf16 lora for LTX.
targeting video, audio, and cross-modal a/v modules resulted in ~200mb, but since my training data had no relevant audio, I targeted only video weights and that's what reduced it to ~100mb.
So to answer your question.... maybe?
going rank 32 would certainly increase file size, but would it improve the understanding? or just increase the capacity to implement what it learned from the training data? I dunno. I think the bigger lever would be a lot of well-curated feets training data, but that wouldn't affect the lora file size.
I really think the biggest problem is LTX having such rudimentary understanding of feet. Not sure this is something that can be corrected with a lora trained on a lil 5090. At least not outside specific contexts like 'footjob'
LTX keep changing feet to hands... how to overcome this :((((((