I compiled a dataset of 50+ video clips (3–15 seconds long) for training, with approximately 50% consisting of close-up front-facing shots. I extracted frames at 1-second intervals and used Gemma to generate descriptions following the Minimax requirements. Finally, I manually reviewed and polished the captions to fix any obvious errors.
The LoRA works well with the fl2v initial frame. Although there are a few examples of putting on a hook in the training set, they are very sparse, so the LoRA usually struggle with it - it’s better to use a frame change or pre-draw it using image editing. Note that there are no examples of strictly top-down or back views.
Occasionally, the LoRA might generate extra hooks hanging on the sides of the head, so tweak the prompt or try a different seed if that happens
Description
Start your promt with
integrated_multimodal_description: [Shot 1] Nosehook.