For FL2VA Model.
I have a lot of different footjob datasets, so I plan to post multiple "types" on this model card. An AIO may come later, but it's still too early to say how realistic that is.
I renamed "v0" as "Type A v1" - it is still a valid, simple footjob lora with no specific motion other than 'toes wrapping around the cock."
TYPE A STRENGTH: 1.0
"Type B" is a toejob with toes gripping tightly, cock partially between hallux and toes, and the girl giving more attention to the head of the cock. This is a much stronger lora than Type A.
TYPE B STRENGTH: 0.4~0.7
0.5 seems good for general use. Above 0.5, the camera will really want to zoom in on the feet. Above ~0.8, artifacts from the training data start to appear.
Technobullshit:
I switched from Diffsynth-Studio to Musubi-Tuner for Type B (only because I am familiar with musubi from Wan). I was having terrible luck training with the fl2va "task" in musubi, would not seem to learn motion much at all. So I switched it to t2va and seem to be getting better motion learning, possibly because it's not pulling first and last frames that are nearly identical, but tbh I have no idea.
I experimented with lower res cache data (192 px) and found it trained quickly but easily overbaked and low res artifacts started appearing in my generations before the movement did. However, this did seem to be a good indicator of how well the dataset and parameters would work for higher res training. This seems like a useful method for relatively quick checks for future training runs.
I also experimented with rank 32 / alpha 32, 32/1, 16/1 and 16/16.
The final published Type B was trained on five 512x512x124 clips, one of which had repeats=6 so basically 10 clips. rank 16 alpha 16, lr 1e-4, trained on the int8_convrot. swapped 32 blocks and took about 15 hours to train to 400 epochs/4000 steps on the 5090.
Description
FAQ
Comments (3)
ref2v plz
Works and looks fantastic, great job.
@icelouse make a variant, squeeze head between big toe >w<