CivArchive
    wushu_action v8.safetensors - v8.0
    NSFW

    runninghub cn:https://www.runninghub.cn/ai-detail/2095217490568761345

    runninghub ai:https://www.runninghub.ai/ai-detail/2093900524172447746

    huggingface:https://huggingface.co/Jojocodex/wushu-action-v7-minimax-h3-fl2va-ref2va-lora

    使用前必读

    1. 触发词 wushu_action 必须放在提示词最前面。所有训练标注均以此词开头,模型将该词与武术动作强绑定。

    2. 该LoRA训练素材全部为人体模型三视图动画,默认容易输出白模+纯黑背景。这是设计特性,不是BUG。

    3. LoRA权重推荐范围:

      • 0.9–1.0:生成三视图人体动作参考素材

      • 0.6–0.8:正常场景出片(推荐起点)

      • 若大量出现白模黑底,降低至0.5–0.6

    提示词写法

    模式A:生成训练风格三视图人体动作参考

    直接复制模板,仅替换招式描述

    wushu_action, 白色无贴图男性人体模型, 同一角色三视图同步并排, 纯黑虚空, 固定全身机位, 动作完全同步, 不是三个战士,【填入招式】, 实时爆发力, 全身发力传导, 无慢动作, 无停顿
    

    模式B:生成真实角色与场景(视频创作推荐)

    保留触发词和动作关键词,把模板里的人体模型、三视图、黑底替换为你需要的角色和场景

    wushu_action,【角色与场景描述】,【填入招式】, 实时爆发力, 全身发力传导, 无慢动作, 无停顿
    

    提示:优先使用训练集中已存在招式词,成功率更高:直拳、高侧踢、疾步冲刺、举盾格挡、长枪突刺、翻滚、连续连招

    H3 生成参数

    • 采样器:euler|调度器:simple

    • 采样步数:25

    • CFG:1.0,请勿大于1

    • 视频帧数:遵守 17n+5 规则(可选:22 / 39 / 73 / 90 / 124)

    • 分辨率:宽高必须为32的倍数,推荐 832×480

    负面提示词

    模糊,扭曲,低画质,画面抖动,多余肢体,手部畸形,形体变形,光影错乱,画面闪烁,静止姿势,动作僵硬
    

    已知限制

    • 训练片段时长仅1.3–3秒。长视频建议分段生成后剪辑拼接。

    • 训练素材均为单人,双人对打属于外推效果,稳定性较差。

    • 本LoRA只学习动作;角色皮肤、服装、面部质感依靠底模或画风LoRA实现。

    ComfyUI部署方法

    1. 将LoRA文件放入 models/loras/ 文件夹

    2. 加载 minimax_h3_fl2va 扩散模型,搭配对应的Qwen3VL文本编码器、H3视频VAE

    3. UNet加载节点后接入LoRA加载器,设置上面推荐的权重

    4. LoRA挂载校验:对比权重0与权重1的出片效果。若无明显差别,需要对safetensors做键名重映射。

    Wushu_Action LoRA Official Usage Guide

    Mandatory Prefix Rule

    The trigger word wushu_actionmust be placed at the very beginning of every prompt. All training data is anchored to this prefix, which strongly binds the model to recognize and generate authentic martial arts movements.

    Model Characteristics

    This LoRA is trained exclusively on 3-view orthographic human animation references (front/side/back). Default outputs tend to feature white textureless mannequins with pure black void backgrounds. This is an intended design feature, not a bug.

    • 0.9 – 1.0: For generating standard 3-view human motion reference assets

    • 0.6 – 0.8: General scene video generation (best starting point for most creations)

    • 0.5 – 0.6: Use if white mannequins or black empty backgrounds appear excessively

    Prompt Writing Templates

    Mode A: 3-View Orthographic Motion Reference (Training Style)

    Use the full template, only replace the martial arts move description:

    wushu_action, white textureless male human mannequin, same character displayed in synchronized 3-view layout, pure black void background, fixed full-body camera, perfectly synchronized movement, no three separate characters, [insert martial arts move], real-time explosive power, full-body force transmission, no slow motion, no pause

    Mode B: Realistic Character & Scene Video (Recommended for Video Production)

    Keep the trigger word and motion keywords; replace mannequin/3-view/black background terms with your custom characters and scenes:

    wushu_action, [character & scene description], [insert martial arts move], real-time explosive power, full-body force transmission, no slow motion, no pause

    High-Success Pre-Trained Motion Vocabulary

    These movements exist in the training dataset and yield the most stable results: Straight punch, high side kick, rapid dash, shield block, spear thrust, roll dodge, continuous combo attacks

    H3 Generation Parameters

    • Sampler: Euler

    • Scheduler: Simple

    • Sampling Steps: 25

    • CFG Scale: 1.0 (do not exceed 1.0)

    • Video Frame Rule: Follow 17n+5 sequence (Recommended: 22 / 39 / 73 / 90 / 124 frames)

    • Resolution: Width & Height must be multiples of 32. Recommended: 832×480

    Negative Prompt

    blurry, distorted, low quality, screen jitter, extra limbs, hand deformity, body distortion, wrong lighting, flickering screen, static pose, stiff movement

    Known Limitations

    • Original training clips are only 1.3–3 seconds long. For long videos, generate segmented clips and edit them together.

    • All training data features single-person movements. Dual-person combat is an extrapolated effect with lower stability.

    • This LoRA only learns motion logic. Character skin, clothing, facial details and rendering styles rely on the base model or style LoRAs.

    ComfyUI Deployment Guide

    1. Place the LoRA file into: models/loras/

    2. Load the minimax_h3_fl2va diffusion model, paired with the matching Qwen3VL text encoder and H3 Video VAE

    3. Connect the LoRA loader node to the UNet node and apply the recommended weight values

    4. LoRA Validation Test: Compare outputs at weight 0 and weight 1. If no visible motion difference exists, remap the safetensors key names manually.

    Description

    FAQ

    LORA
    MiniMax H3

    Details

    Downloads
    222
    Platform
    CivitAI
    Platform Status
    Available
    Created
    10/5/2026
    Updated
    10/7/2026
    Deleted
    -

    Files

    Minimaxh3多参考双采打斗工作流 .json

    wushu_h3_v8_final_fp16.safetensors

    Mirrors