What actually changes
You add one file and drop the step count from 20 to 8. Nothing else in your setup moves. Measured here at 960x544 and 124 frames: 113 seconds becomes 53. At 4 steps it is 34 seconds, which is what I use while I am still hunting a prompt.
My rule: 4 steps while you are searching, 8 for the one you keep. When you do not yet know whether a prompt works, you do not need the best version of a bad idea. You need to find out quickly that it is a bad idea.
What you need
minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors →
ComfyUI/models/loras/(this is the one that makes it fast)minimax_h3_fl2va_pruned_int8_convrot.safetensors →
ComfyUI/models/diffusion_models/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors →
ComfyUI/models/text_encoders/minimax_h3_video_vae_fp16.safetensors →
ComfyUI/models/vae/minimax_h3_audio_vae_fp32.safetensors →
ComfyUI/models/vae/
The folder is the part people get wrong. A correct file in the wrong folder looks exactly like a missing model, and ComfyUI will not tell you which one it is.
Two things that still catch people
The type box still has to say minimax. Changing the step count does not change that, and if it is wrong nothing runs.
You need both decoders, one for the picture and one for the sound. Leaving either out fails in a way that does not obviously say so.
I wrote up every setting, what each one does and the mistakes that cost me time, in the full written guide. Free, and there is no card.
There is also a paid course that goes further with this model if you want it. Either way the workflow here is the whole thing, not a trimmed version.
Description
FAQ
Comments (1)
I rushing to test this out right now!😲
Awesome work as always! 🙏👍

