For FL2VA Model.
๐ Multiple Footjob Varieties Available! ๐ฃ
โโโโโโโโโโโฌโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Version โ Strength โ Description โ
โโโโโโโโโโโผโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Type A โ 1.0 โ Simple footjob, toes gripping shaft on both sides โ
โ Type B โ 0.4~0.7 โ Toes grip on both sides, more 'head' action โ
โ Type C โ 0.6~0.8 โ One foot stationary, one foot moving โ
โโโโโโโโโโโดโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโTechnobullshit:
Type C notes...
Learning more and more about H3 training but still not totally nailed down.
I am still struggling with the tendency for it to zoom in on the feet or make the penis extra huge. I assume this is because the training data contains a lot of close-ups which it is trying to reproduce. In an effort to correct this I trained Type C with a reduced max_timestep, 700. Trying to apply the same kind of high/low noise logic I learned from Wan 22. 700 is probably still too high but I didn't want to eliminate all high noise movement training. This deserves further experimentation. I also increased min_timestep a little bit, to 100, in an effort to reduce the compression-artifacteyness seen at higher strengths. It maybe worked?
Type B notes...
I switched from Diffsynth-Studio to Musubi-Tuner for Type B (only because I am familiar with musubi from Wan). I was having terrible luck training with the fl2va "task" in musubi, would not seem to learn motion much at all. So I switched it to t2va and seem to be getting better motion learning, possibly because it's not pulling first and last frames that are nearly identical, but tbh I have no idea.
I experimented with lower res cache data (192 px) and found it trained quickly but easily overbaked and low res artifacts started appearing in my generations before the movement did. However, this did seem to be a good indicator of how well the dataset and parameters would work for higher res training. This seems like a useful method for relatively quick checks for future training runs.
I also experimented with rank 32 / alpha 32, 32/1, 16/1 and 16/16.
The final published Type B was trained on five 512x512x124 clips, one of which had repeats=6 so basically 10 clips. rank 16 alpha 16, lr 1e-4, trained on the int8_convrot. swapped 32 blocks and took about 15 hours to train to 400 epochs/4000 steps on the 5090.
Description
FAQ
Comments (8)
Do you think it's too much for Minimax H3 to train multiple concepts into one lora? What you're doing works and is cool, but I was just curious about limitations.
From what I've seen so far, it's definitely possible to combine.
But being barely two weeks in, I'm still figuring out H3's quirks both in training and inference, and it's easier to do that with one clean concept at a time.
I have 2 or 3 more distinct footjobs I want to train individually before I start trying to combine them, but a future all-in-one is very likely.
Fantastic work!!
Great work, are you planning maybe doing a normal footjob lora where both feet are straight and moving paralell up and down, very normal footjob, because in my gens she keeps moving the feet like she is riding a bike or something
Yes definitely. The dataset will likely overlap with my 'shoejob' lora, which I 100% guarantee is on the way ๐ And I totally know what you mean about the bike pedaling footjobs in vanilla H3
TypeC is the dream thanks for making it!
Could you make a version that can recognize stockings?
unfortunately i can't find your checkpoint of h3 :(