VDN is a new attention system for MiniMax H3. It makes video generation much more efficient by keeping normal attention where it matters for nearby frames, while using a much cheaper linear attention mechanism for long-range information across the video.
INT8 ConvRot VDN stage checkpoint prepared for ComfyUI-VDN-H3-24GB node to use on custom PCs with 24Gb GPU.
In simple words, this adapter (model) and my ComfyUI node allow you to generate videos in 8 steps without turbo LORAs and get a quality close to the quality of the original model in 20 steps.
If you ask is it faster? I will definitely say yes, it is faster than generating 20 steps on the original model. But if you compare the generation time with generation using turbo LORAs, it will be a little slower, but the waiting time is worth it if you need quality. You can see the test results below.
How to use it:
download archive "stage-dmd-step-250-int8_convrot_comfyui" unzip it and put it in a directory ComfyUI/models/vdn/ full way should be ComfyUI/models/vdn/stage-dmd-step-250-int8_convrot_comfyui/. !!!do not change the names of the models!!!
Same model on HF: speach1sdef178/VDN-H3-INT8-ConvRot-ComfyUI · Hugging Face
The final local structure must be:

Required ComfyUI node https://github.com/Speach1sdef178/ComfyUI-VDN-H3-24GB
The node has been added to the ComfyUI Manager
24 GB-oriented MiniMax-H3 VDN runtime for ComfyUI. Version 1.1.0 keeps the validated v49 AutoMemory/AutoLongCache policy and adds three narrowly scoped correctness and memory-safety fixes tested on an RTX 3090 24 GB.
Example WF in the attached files

License: This model is a derivative of MiniMax H3 and is subject to the MiniMax H3 Community License Agreement.
Description
VDN is a new attention system for MiniMax H3. It makes video generation much more efficient by keeping normal attention where it matters for nearby frames, while using a much cheaper linear attention mechanism for long-range information across the video.
INT8 ConvRot VDN stage checkpoint prepared for ComfyUI-VDN-H3-24GB node to use on custom PCs with 24Gb GPU.
How to use it:
download archive "stage-dmd-step-250-int8_convrot_comfyui" unzip it and put it in a directory ComfyUI/models/vdn/ full way should be ComfyUI/models/vdn/stage-dmd-step-250-int8_convrot_comfyui/. !!!do not change the names of the models!!!
same model on HF: speach1sdef178/VDN-H3-INT8-ConvRot-ComfyUI · Hugging Face
The final local structure must be:
ComfyUI/
└─ models/
└─ vdn/
└─ stage-dmd-step-250-int8_convrot_comfyui/
├─ model_spec.json
├─ linear_branch/
│ └─ model_int8_convrot_comfyui.safetensors
└─ adapters/
├─ default/
│ ├─ adapter_config.json
│ └─ adapter_model.safetensors
└─ turbo/
├─ adapter_config.json
└─ adapter_model.safetensors
Required ComfyUI node https://github.com/Speach1sdef178/ComfyUI-VDN-H3-24GB
24 GB-oriented MiniMax-H3 VDN runtime for ComfyUI. Version 1.1.0 keeps the validated v49 AutoMemory/AutoLongCache policy and adds three narrowly scoped correctness and memory-safety fixes tested on an RTX 3090 24 GB.
Example WF in the attached files
License: This model is a derivative of MiniMax H3 and is subject to the MiniMax H3 Community License Agreement.
FAQ
Comments (22)
Attention!!! This particular node is needed!
Thank you ❤️
"!!!do not change the names of the models!!!" Yeah don't count on CivitAI for that 😂
Both files renamed to "int8ConvrotVDNH324GB_v10" by CivitAI.
@GlowingGuardianGirl oh, really? ha ha ha thanks! ))
@speach1sdef178 Yeah sadly every file is renamed, I have no idea why they keep doing this. @CivitaiOfficial @JustMaier <<< STOP RENAMING FILES!
@GlowingGuardianGirl yes it's not good
@GlowingGuardianGirl I downloaded my archive and the files in the archive are not renamed, the names are saved and each model is in its own directory. But if that's not the case, just download the model from https://huggingface.co/speach1sdef178/VDN-H3-INT8-ConvRot-ComfyUI
A small hint: if you enclose some text in triple backticks, 3x(`), your directory structure should be preserved, like this:
ComfyUI/
└─ models/
└─ vdn/
└─ stage-dmd-step-250-int8_convrot_comfyui/
├─ model_spec.json
├─ linear_branch/
│ └─ model_int8_convrot_comfyui.safetensors
└─ adapters/
├─ default/
│ ├─ adapter_config.json
│ └─ adapter_model.safetensors
└─ turbo/
├─ adapter_config.json
└─ adapter_model.safetensors(Copied from HF)
nice work here can i may integrate it into me newest project https://civitai.red/models/2935896/furry-enhancer-studio ?
@freek22 hello :) didn't see you a long time! glad to hear you :) of course, you can use it, it's for all )))
wait for real spedup . on my 5090 for 10 sec clip in normal resolution H3 minimu 6 min.
LTX25 after upscale 20 sec clip 1.5 min . with loras good prompt not worse resutls. Is any chance to make H3 usable for longer films?
longer clip? how much longer? I think 20 sec clip it's more than enough for the start and then generate the next clip from the last frame of the first generation. no? :)
@speach1sdef178 I made ususalle over 10 min videos. So count the clips needed.
@3dasdman I think this question for the developer, I think. I am not the author of H3 and not from their team, I will not be able to tell you when a model will appear capable of generating a 1-hour clip (conditionally) and whether we will be able to run this model on our equipment or not?! I usually generate 5-10 seconds of video and if I need to, I will glue them together in the way that is currently available. What I have shared is just a technology that allows you to return details to generation when we using quantized models.
The audio sounds distorted/metallic (similar to the 4step turbo loras). Which is also present in the previews posted. Any idea how to fix it? I tried a couple different shift values and it doesn't seem to get better.
@SweetAyanna sound it's a trouble of the model itself. Right to say it's a trouble low steps. But vdn it's not about sound, it's about pixels
@speach1sdef178 Actual answer... You can generate audio fast with a whole render at 32x32 high steps, then use that audio output as a reference input for the video. By the time you do all that, I'm not sure it's much faster.
@speach1sdef178 Most of the 8 step turbo loras do not destroy the whole audio like that. And personally, I do not see a visual improvement over lets say ema600 or native 8 step lightx2v turbo loras. If a method guts audio in a multimodal model while arguably not providing a noticable benefit, maybe the method is just not ready yet.
@SweetAyanna I don't force you to use this method if you don't like it. I just shared the result of the work, and I use it because I see a huge difference. everyone chooses what they like and feel comfortable with
and regarding the sound in the attached videos, I used audio references, but I didn't describe them in the prompt because that wasn't the goal at all. so there is no need to evaluate the sound quality here. VDN is not designed to enhance sound
@speach1sdef178 No worries. I just noticed the visual was good and wondered if there is a way to improve the audio. Which is impressive for 3 steps either way.
@SweetAyanna I think there are options, but I haven't started working on it yet. I am currently working on a large project for H3, and I will try to work with Audio after completing the current project.