! Image to Video generations should be fixed now !
V7.5 Update
Image to Video generations should be fixed now.
***- PLEASE leave some kind of feedback after trying -
Base Information
Now actually the fastest HQ audio + video generation workflow. See for yourself in the examples. I can currently create these 8s, 1 megapixel videos in 3.5 minutes which is INSANE if you ask me, while keeping good audio!
I'm still experimenting with different scenarios, tweaking for the ultimate best balance of quality vs speed. There might be more updates very soon.
WHY 2 MODELS AND REFINEMENT AT ALL??
Yes, valid question. My answer: creating high resolution videos with the Fast H3 model takes much longer than this workflow, and the quality of the short Taomate turbo workflow is outright bad with low steps. The Solution: We use the power of the audio quality and base movement of the Fast H3 model while then using the ultra-rapid speed of the Taomate 3 step turbo lora to refine and enhance the video to High Quality!
PLEASE let me know how this works for you!
This is my very first attempt at making a FAST workflow that runs on my 16gb + 32gb System RAM Setup and I'm looking forward to see if you guys have the same results like me.
Please try it out and let me know what you think!
INSTRUCTIONS:
Use the green nodes to configure your resolution, aspect ratio & duration
Keep the aspect ratio the same for the first pass and second pass resolution selector
Add your desired LoRAs in the purple node
Enable or Bypass your desired 1st Frame or Last Frame input Image Groups on the left side of the workflow
Adjust to your desired VAE's, text encoders that you might already have (I'm using a new KJ video vae that is smaller)
Don't forget to adjust the filename and path in the Save Video node to your prefered syntax and location (the one in the workflow creates this type of filename syntax with MP, date and time of generation: 1MP_Video_210926-150209_24fps_00001_.mp4)
MODELS
Base FLF2VA / T2V model: https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors
Text Encoder: https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/text_encoders/qwen3vl_32b_minimax_h3_int8_convrot.safetensors
Audio VAE: https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/vae/minimax_h3_audio_vae_fp32.safetensors
Taomate 3 Step Turbo Lora: https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/blob/main/experimental/minimax_h3_taomate_fl2va_3step_ema_comfyui.safetensors
H3 Latent Upscaler: https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler
Custom Nodes Links
GENERATION SPEED:
On 5070Ti with 32gb System RAM:
8 seconds - 1 MP resolution: ~3:30 minutes
8 seconds - 0.5 MP resolution: ~1:45 minutes
10 seconds - 1 MP resolution: ~4:30 minutes
10 seconds - 0.7 MP resolution: ~3:10 minutes
Description
Adjusted 2nd Pass Lora strength to 1.0.
Set 0.4MP as new base resolution for 1st pass because it creates much more details while taking almost the same amount of time.
FAQ
Comments (13)
Sadly my RTX 4060 Ti 16GB can’t run nodes #181 and #187 properly, but the workflow still works okay when I launch ComfyUI with --disable-pinned-memory --vram-headroom 1.
I’m quite new to ComfyUI, MiniMax H3, and local open-source video models, so I don’t have a lot of experience yet. That said, I also think the idea others mentioned about making a Ref2V version of this workflow is really important.I’m quite new to ComfyUI, MiniMax H3, and local open-source video models, so I don’t have a lot of experience yet. That said, I also think the idea others mentioned about making a Ref2V version of this workflow is really important.
I also rearranged and localized the workflow for my own use (screenshot attached). I also rearranged and localized the workflow for my own use (screenshot attached: https://imgur.com/a/uHfxwIy).
If you ever need someone to help tidy up or re-layout already-stable workflows, I’m happy to offer some free help for now. In return, I’d just appreciate the chance to chat more about technical topics in the future — after all, I only recently started with ComfyUI (well my questions so far have been answered by DeepSeek and Grok most time).If you ever need someone to help tidy up or re-layout already-stable workflows, I’m happy to offer some free help for now. In return, I’d just appreciate the chance to chat more about technical topics in the future — after all, I only recently started with ComfyUI (well my questions so far have been answered by DeepSeek and Grok most time).
Thanks for sharing this workflow!Thanks for sharing this workflow!
Hey Thanks! Your workflow version really looks tidied up, but I would keep some things like the final resolution selector in a main control area on the left side to adjust duration, size, etc. in one place. Also I would reposition the Lora Loader because you will need to see the names of the selected Loras, and many Loras have very long names, also the node itself will increase in height with each Lora, so it's best to have some empty space beneath it to not break the UI after too many are loaded. I also updated my workflow a few times since the version you have and both those nodes are not longer in use in my workflows. SolAttn is deprecated by now and instead of MemEffSageAttn we use the ComfyKitchenAttention node instead.
Thanks for sharing your version! I will keep you in mind when it comes to WF optimizations :)
@zkyt yeah I just see the v7 version, sorry for my lately comment, I will try again these day ;>
Thanks for the workflow. Question, how come my final generated video looks different then the reference image and even the first preview clip? Once it goes through the latent upscaler it changes everything about the video, person changes, clothes change and background as well. How do I make the final video look like the image and the first preview video?
I just did some research and I think you brought up a point that makes this workflow invalid for I2V generations. Because the 2nd pass of the workflow currently doesn't receive the input image, and I guess you are referencing the starting image somehow within your prompt, it does not understand what is meant and will therefore create weird stuff. I will take a look at this further. In my I2V tests I only used prompts that were mentioning the subject in the image directly like this "the woman suddenly turns around and starts running" and it worked without mentioning the first frame image at all. I guess that's why my tests worked and you run into problems. So it's a workflow+prompt issue is what I think. Thanks for bringing this up!
connect your first frame image (from the resize node output underneath the load image node) to the 2nd conditioning node in the upscale section.
+ reduce denoise strength in the upscale section helps too (even without connecting the image but it'll be less precise)
@Crezjcm I can't believe I didn't see that when troubleshooting... Thank you very much for the hint! I have also connected the Last_Frame image input correctly now.
@kbradpatsnation647 I fixed it.
The very first 2 people who downloaded it might have the wrong file, because the website didn't wait for me to actively publish. Please re-download if the workflow is missing the image connection in the 2nd pass. (v7.5)
Hi, wanted to say thanks for making this workflow. I have a few questions and comments.
Could you show us where the workflow may be missing the image connection?
I have a few suggestions to make to speed things up. It's great using comfy kitchen attention but I somehow got a boost of speed when I paired Spectrum node and Sage Attention KJ nodes (compile set to true) together towards the end of the model chain. Maybe it's using VRAM efficiently but it speeds up the upscaler at the end.
@daltonibarra761431 Hey, thank you! The missing connection was on the right-top side of the workflow, missing the first_frame and last_frame connection, it hit 1 or 2 people at maximum who downloaded it about an hour ago.
I'm not really convinced that Spectrum is worth it for the 2nd pass, because it helps skip steps and the refinement already only has 3 steps in total. Regarding the Sage Attention node: for me in my tests it didn't make any change in the generation speed, and because the standard workflow from Fast H3 uses Comfy Kitchen, I re-use it to keep it more consistent. Please let me know of your findings if you keep testing it!
Thank you!!! I was so confused but I managed to fix it manually without realizing. I think you are right about the second pass - it actually makes the audio wonky or it could be the strength of the 3 step LORA - I believe it recommended a strength of 0.75 but I'm not 100% sure. I was doing a bit of tweaking today after work but I will definitely have some updates for you tonight or tomorrow.
@daltonibarra761431 Sounds great! Also try lowering the denoise from 0.35 to 0.25, if you want to test more detail.