This workflow has been replaced with my new MINIMAX SEED HUNTER WORKFLOW: https://civarchive.com/models/2881362/minimax-seed-hunter-workflow-optimized-fast-latent-upscaler-speedups
Please use that instead.
Watch the tutorial video on how to build an anime-style short using this workflow:
A no-nonsense high level T2V / I2V / FFLF / REF2V workflow for Minimax H3 with lots of options & togglable quality of life features.
Toggle between the [T2V / I2V / FFLF] model & the [REF2VA] model easily
Added Forced Custom Audio: Lipsyncing to custom dialogue/music is now easy for T2V/I2V/REF2VA!
Temporal Upsampler 2nd Pass Option! [https://github.com/matlowai/ComfyUI-MAINodes]
New Image Loader allows cropping/setting max megapixels directly inside node: https://github.com/obvpm/comfyui-obvpm (Must install via unzipping to custom_nodes folder or through comfy-manager's Install via Git option, it's too new to be on the comfy-manager's index)
Kijai's Preview Override (put https://huggingface.co/Kijai/MiniMax-H3-TAE/blob/main/vae_approx/taeh3.safetensors in /vae_approx/ and set as the custom vae for non-pixelated previews)
Speedup: Sage-Attn + New Sol-Attn [https://github.com/kijai/ComfyUI-SolAttn_triton]
Speedup node: EasyCache / Spectrum [https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3]
Turbo LoRA (t2v/i2v but works with ref2va if you don't mind [some] audio degradation)
New (8/11/26) Lightx2v 1.0 release [https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main]
Film VFI Frame Interpolation (24 fps -> 48 fps)
Easily togglable reference fields: 4 pictures, 2 audio, 1 video
Use this workflow if you:
Are an AI Filmmaker who values character/scene consistency
Want to be able to use FFLF effectively, seamlessly extend videos, or use character reference sheets.
Want to edit videos, or copy & use motion from reference videos
Want fine-tuned control over your shots visuals and audio.
If you wish to generate simple one-off video clips like Will Smith eating spaghetti, the t2v/i2v model that you can toggle to in this workflow will do that for you.
Kijai's int8 convrot video vae: https://huggingface.co/Kijai/MiniMax-H3-experimental/blob/main/minimax_h3_video_vae_int8_convrot.safetensors
UPDATE YOUR COMFY CUDA VERSION TO 13.0. If you start comfy and see cu130:
[INFO] pytorch version: 2.13.0+cu130Then you are good to go! But if you are using cu126, then ALL your gens with the best version of Minimax's model (INT8 convrot) will be 2x slower than they should be due to inefficient comfy-kitchen operations! Keep in mind if you update your cuda, you will need to reinstall Sage Attention! Download the correct wheel from here https://wildminder.github.io/AI-windows-whl/
ALL MODEL / NODE LINKS ARE IN THE WORKFLOW NOTES OR EASILY INSTALLABLE THROUGH COMFY-UI MANAGER.
If you appreciate what I'm doing, please consider following/subscribing on my patreon (free) which gets access to all my work early.
https://www.patreon.com/cw/foxfuressence
Also my youtube where I make AI filmmaking tutorials plz & thx:
Description
Edit: 8/16/26: Fixed bug where Audio 1 wasn't plugged in to H3 REF2VA node.
Adds Comfy-Kitchen Attention Speed Up
Adds Temporal Upsampling De-Rope (fixes fast motion blur/smudge) [https://github.com/matlowai/ComfyUI-MAINodes]
Replaces RIFE with built in comfy node Film VFI
Adds Model Shift node
This is a pretty large update. The Temporal Upsampling 2nd Pass option improves action scenes drastically.
FAQ
Comments (98)
Thanks for making this great workflow. i used to use it from the beginning,. it's greatly designed. and i would have the following suggestion. actually i modified it to add the following for my routine use. and i guess you might also interested.
1. Add an LTX 2.5 upscale 3 steps for FL2V with original prompt input before RFIE, it can efficiently upscale the output quickly. The overall time for high resolution output can be reduced,
2. Addition of optional (default off, it's too slow) SeedVR2 Video upscale (have modified INT8 node). This is very slow, but just in case someone want to output high resolution video without much loss of quality.
3. Addition of clean VRAM step and RTX upscale after RFIE, So the resolution can "Quickly" increase, this is fast, but nvidia card only. not for ROCm.
Anyway, thanks for updating the workflow to 2.0.
And i have 1 question. can F2V do lip sync with audio input? it doesn't seems work... but ref2va is ok.
I have thoroughly tested LTX 2.5 upscale and it is extremely bad for several reasons: 1. It is prone to the same problems LTX has (any high motion shots turn into smudge/screen tearing), 2. it requires an entire other model load which cuts down on gen speed overall, and 3. the prompts for minimax look so different than the prompts for LTX that you can't even really gain any guidance benefit.
I am also not a fan of SeedVR2's video upscale, having tested it thoroughly. It's jittery and even with large context sizes (batch size I think they call it?) it's so clearly unintelligent about what it thinks blurry objects are that the extra detail makes the output worse. Same with RTX upscale.
Ultimately if I add an upscale it will be what I am currently testing: Upscaling the latent directly after 1st pass finishes, and use Minimax to upscale with just a few extra steps.
And yea, using the Force Custom Audio option you can make both models lip sync really easily. Just toss the line of dialogue into the prompt between <d></d> tags, enable Audio 1, load your line of dialogue, and switch that Force Custom Audio option on. It overrides all other audio tho, so be aware of that.
One more suggestion. this is also one of my modification from your workflow.
Since the turbo LoRA for FL2VA and REF2VA will be two different LoRA. when i switch from FL2VA to REF2VA, i usually forget to switch the LoRA too.. i suggest to load the LoRA before the Anyswitch, so you can make two different LoRA node for different model.
I thought about that but ultimately didn't want to force the issue for anyone. Most people aren't swapping between models all that much, it just felt a little too hand-holdy for my taste.
It looks good. But all the necessary parameters are scattered in random places. And... I can't find the node to set the animation time for FFL2VA.
And it’s not clear whether FFL is a typo or is it different from FL?
It's all in the green group node area labeled "Video / Prompt Settings." The node is "Video Length (seconds)". I am aware this is a fairly complicated looking workflow at first. It isn't randomly scattered at all. Each color-coordinated group is one theme: Red on the left is optional functions, green is prompt/video settings, and yellow are user inputs/references. Everything is toggleable and flexible, so there's no avoiding SOME initial confusion. And yes, FFL is just a typo for FFLF, thank you for pointing that out.
Is it normal for reference to video while using a video reference to take SOOOO long to render? I used your last updated workflow so not this new one but I have an rtx 5090 and at 0.7mp trying to copy a 11sec video took me 25 minutes I believe 😭
Are you using the new Temporal Upsampler? If so then yeah, that adds a bunch of time to a gen because it's literally going back through it and regenerating the smudged/low-quality parts.
If not, V2V editing just straight up takes a long time with this model. If you're referencing a long video (over 24 frames) then yeah, it can take that long. Video edits of 96+ frames or more, easily.
@foxydits so there’s no way to really go around that then? Looking forward to trying your latest workflow as I have really liked your first 2, because the regular I2V is really fast
@Noob_models t2v, i2v is fast, ref is slow, v2v is slowest. it's the nature of the beast.
try skipping some frames in reference video. this helped me a lot.
In version 1.8, REF2VA worked perfectly with reference audio—my character matched the reference pitch precisely. However, in 2.0, it seems to lose that pitch matching.
Also, could you consider bringing back RIFE Frame Interpolation in a future update?
The updates to my workflow don't control/change how the model handles audio, so that one's not on me.
Is it because Film VFI is slower? I switched because it's a) built into comfy, so ppl don't have to get more nodes, and b) it's higher quality (RIFE is faster for a reason)
Film VFI is slower, if i changed the fps to 60 ,audio will out of sync, but RIFE doesn't had this issues and much faster
@addison8406 Not too had to switch it back to rife yourself. Just load the node and replace the links with the ones connected to film VFI.
@foxydits thanks for your advice, i switch back to rife now, and i just found out a weird issue, if i put a reference audio to <Audio 1> -- not working, character won't matched the reference pitch, it will gen a new voice of my character,
force <Audio 1> still working, character speaking the same sentence from reference audio
if i put a reference audio to <Audio 2>, worked perfectly~!!!
er_sde/beta breaks reference audio,the audio is glitchy, switch to res_multistep/simple or sgm_uniform the reference audio will be more stable and clear
thanks for your effort, i really love this workflow
@addison8406 You have the version I released with audio 1 not hooked up, I fixed it early yesterday and put a note up about it on version info. Sorry about that.
er_sde/beta produces perfect audio for me, and is what pretty much everyone in the Banodoco channel was suggesting too. Sorry it's not working for you but it certainly does for most of us! Protip: If you use turbo LoRAs and the ref2va model, try plugging the fl2va model into the ref2va node. They function exactly the same when configured this way but fl2va is faster and the turbo lora actually works!
@foxydits if i plugging the fl2va model into the ref2va node, which turbo LoRAs should i use?
minimax_h3_fl2v_lightx2v_turbo_8step? or minimax_h3_ref2v_lightx2v_turbo_4step?
@addison8406 use the fl2va turbo lora.
So... amazing! This workflow is incredible.
I have two feature requests if possible:
Multi-shot support — the ability to generate multiple 10-second clips (for example, 6 clips) and automatically concatenate them into one continuous video.
Additional prompt support — an option to add extra prompts/instructions for each shot or clip, so the actions and transitions can be controlled more precisely.
This workflow is already amazing, and these features would make it even more powerful!
Hi! Thanks for the glowing review!
Context window multi-shot concept is not really compatible with this workflow as far as I'm aware. Unless you're going for longer than 15 second continuous shots, there's really no purpose for it either. This is an AI Filmmaking workflow, so the assumption is that you have experience with a good video editor like Davinci Resolve and have no problems connecting the clips into a finished piece. All creators owe it to themselves to acquaint themselves with a video editing software. The quality of your gens goes up so much when you know basic editing, as shown in my video posted on this workflow.
I do like the idea of adopting a prompt-helper node, something with buttons to auto-write out code-blocks that are common for the ref2va model. "subject_definitions: <subject 1> is the character in ..." stuff like that.
very nice bro. i will definitely try this
Get Error when Sage-Attention is enabled
# ComfyUI Error Report ## Error Details - Node ID: 125 - Node Type: SamplerCustomAdvanced - Exception Type: AttributeError - Exception Message: AttributeError: 'NoneType' object has no attribute 'qk_int8_sv_f8_accum_f32_fuse_v_scale_attn_inst_buf'
Amazing work, as always! Any chance we'll get a "continue video" option (like in your LTX workflows) for the FL2VA model?
I think that's doable now with the h3 Reference Guide option isn't it? I will look into it today.
Fails due to insufficient GPU memory (VRAM) at the Temporal De-Rope Sampling stage on an RTX 5070 (12GB), 15 sec, 480x864, 8 steps. How to fix?
By the way , how do you change the duration? , I found length 124. but i cannot change it. I only got 5 sec video.
@zxyyxrxx3303 Video Length (seconds) node in Video / Prompt Settings group, where is the prompt.
@Spaceknyte thanks, i found it. good work sir. REF work flow very much Ok so far, But when I switch to FLV workflow, only noise video i got. Any advise sir?
Deroping takes a really high amount of VRAM, it effectively lengthens your video by creating freeze frames of high speed moments. a 15 second video to start is just pushing your card too far. There is an option in the deroping nodes set to "Balanced", you could switch that to "Performance" or whatever the economy option is called. But even I with 24 GB VRAM and 64GB system RAM only derope videos that are 8 seconds or less, otherwise it just takes too long. Minimax is a memory hog.
@zxyyxrxx3303 I've only used REF2V in this workflow; the standard FL2V from the repository suits me just fine.
@foxydits thanks, I found the error . it is my fl2va model is corrupt , I use another bf16 and it is ok, I will download prune int8 . thanks sir
GUYS USE REDCRAFT TRUST ME
It doesn't follow the prompt well and produces a blurry image.
Since my comment got a lot of dislikes, I'm wondering: what do you all use?
Worked pretty well with my RTX 3060. The detailed explanations were very helpful too. Thanks for sharing this.
Is there a reason why you don't use the motion context nodes in such an advanced workflow? They should make it even easier to put the cuts together, no?
As a "Filmmaker" workflow, I expect everyone using it to have a basic acquaintance with video editors such as Davinci Resolve. This is a production tool more than anything, where you take your good clips and compose them together in your video editor for editing and polishing. That said, almost no single shot in cinema goes beyond 15 seconds long, it's just not that common. Cuts happen typically every 3-10 seconds. There is simply no need for a workflow like this to have motion context to batch from gen to gen. That's my view, anyway.
@foxydits Thank you very much! Really appreciate your answer. How about scenes with more movement (A dancing couple) or if someone sits cross-leged on a carpet at a specific position?
The later should be no problem if I tell the ref2va model to use the last frame as keyframe, right?
But for the couple in motion a cut might completely change the direction of their dance and the velocity. Would you solve this via prompting, cuts to specific angles like a closeup and Davinci effects so one does not recognize the change in direction/velocity of the new clip?
@dariosmanaris299 Most of the time for shot angle switches and continued motion, my prompts will look like:detail_description:
[Shot 1]: A close up profile side view shot of the couple dancing. The man begins to twirl the woman.
[Shot 2]: A top down bird's eye view of the woman being twirled.
[Shot 3]: A behind the shoulder shot looking at the woman as she finishes her flourish and comes back into the man's arms giggling.
That would be something like a 8-12 second shot that has 3 cuts all within the window that covers the movement. The next gen would use a previous screengrab to establish shot context. Usually like:
`retention_analysis:
<Picture 3> weak_reference is purely a reference for character placement.`
Hope that all makes sense.
Does this work on my computer it says I have 16 of ram I downloaded a file called diffusion_pytorch_model-00004-of-00014 but when I try to open it windows says it does not recognize. What am I doing wrong? I have a red laptop too but I don't know where the plug is.
V2.0 frame interpolation seems a lot slower than RIFE, is that expected?
Yes, it is higher quality though. People can always switch back to RIFE if that's a problem. It's a pretty easy swap.
Thanks for the great work! I am using the 2.0 version and I don't know why the video generated was without audio even I added related prompt. Any idea?
Most likely you downloaded the wf while it still had the missing link connecting Audio 1 to the ref spot. Redownload the workflow. You were just one of the unlucky ones who got I hadn't noticed / fixed it yet. Sorry.
@foxydits Looks like it is not working even I re-download the wf. So you mentioned Audio 1, does it means I have to upload an audio to get it working?
@ericlhm548 I'm confused by the question.. yes, you do have to load an audio file into Audio 1 if you want to use that audio file as a reference in your gen. Like if you load 15 seconds of someone's voice into Audio 1, and it's enabled, in your prompt you would then say:
"<Subject 1> says using <Audio 1> as the vocal-timbre reference: <d>[English]Hey everybody!</d>". You're supposed to also mention it in retention_analysis too, but I think it'll work with just that code alone.
@foxydits Actually what I want is not using the Audio1 but the audio generated by the model itself. It looks like even I use the prompt properly, the video generated got no audio. So I wonder what the problem is. I tried to use the same prompt calling the same model in a very simple wf, it can have audio. So I think there is issue / misconfiguraiton somewhere in the wf .
I tried to form another node and able to export the audio file separately. I think that's some problem on my VHS Video Combine. Thanks for the good work. I think your wf is really GREAT! Can't stop to use it.
Hey, i'm using your workflow for a few days now but i have a problem getting the audio right.
i feel like most of my outputs have clipping audio.
do you maybe have an idea why?
edit: found out why, i needed to increase the steps from 8 to 12 now it's okay
Clipping audio? I'm not sure about that. Though I will say when I released v2.0 there was a window where Audio 1 wasn't hooked up properly. It's supposed to plug into the ref2va node as "ref_audio_0" as well as go to one other location. If you check and it's only plugged into that one other location, but not ref_audio_0, you can either fix it yourself or redownload the workflow. LMK if that was the case.
I have a question. Are the newer versions of your workflows an upgrade to the older ones? Or is each version designed for a specific thing? With the newest one I noticed some different default settings when opening the workflow for the first time.
Each version represents the best possible setup I know of currently at the time of release. The defaults like sampler/scheduler/LoRA/nodes all change to represent what I currently use for all my filmmaking endeavors. 1.6 might have some outdated settings at this point, 1.8 is probably okay, but 2.0 is pretty modern. I plan to remake the entire workflow soon since new tools have come out that can drastically simplify the workflow process.
@foxydits Eagerly Waiting,
Hi. Super nice workflow.. but why it is limited to only 4 images, 2 audio and 1 video reference? how to extend the reference count? because minimax can take much more than this
Space, simply put. I expect users of a workflow like this to know how to copy/paste/connect the links to add what they need. I'm rebuilding the workflow though to be more modular, making adding/removing pieces more easy.
Shouldn't the LoRA Loader be placed before optimization nodes like SageAttention? An LLM analysis suggested that model weights need to be patched by LoRA first before applying execution hooks.
Thanks for sharing the workflow!
I've never personally seen evidence that the order matters there. My rule of thumb has always been that model patches/adjustments should come first, and then LoRAs/shifts come right before sampling. However I have seen workflows linked in both ways that performed exactly the same.
it doesn't matter.
ok, thanks.
after update ComfyUI_RH_MinMaxH3 (H3 Jerk Oracle node):The value 24 for H3 Jerk Oracle (profile / window / hold map)'s abstain_below is above the maximum 10.
so what value should I set now?
This workflow is an absolute maze of experimental patches, redundant switches, and over-engineered hacks that make it a headache
1. it's a brand new model, most of the highest level nodes are "experimental patches"
2. none of the switches are redundant
3. you're redundant for saying the same thing twice ("over-enginered hacks").
@foxydits cry about it big boy
its not a good workflow. its redundant.
@ZIngoFellowGrungitz cough 15.9k downloads cough yeah.. redundant. I guess in the sense that my new workflow is quite nice. man keep projecting your anger, it's only amusement to me.
@foxydits no anger here brother. "new " workflow. the old ones redundant is it not eh?
@foxydits we laugh at you on reddit
https://www.reddit.com/r/StableDiffusion/comments/1vtwtyw/sparse_attention_for_h3_minimax_enjoy_up_to_25x/ This might be interesting to explore/add to the workflow with a future update if it works. Seems to work with CK attention.
Im super confused, how much do these destroy quality? Should i run kitche/solattn/spectrum/sage-attention? if yes all together or just one?
@Gooodis I haven't tested it extensively or anything, but I used it in Plague_Kind's workflow and couldn't see any degradation in quality. I only used CK attention and SLA attention with the 1.1 version of the 4 step lightx2v turbo lora (used it with 8 steps). It was a bit faster and it looked good.
I wish I could figure out how to put the SLA node into this workflow. The instructions say "THE NODE MUST BE LAST IN THE CHAIN, DIRECTLY ATTACHED TO THE GUIDER AND SCHEDULER". But I don't even see a guider? Either I'm blind or it's not there. Anyway I like this workflow a lot better than Plague_Kind's.
This workflow is awesome. It's all over the place, has so much stuff that it's confusing/intimidating, and isn't really arranged in a way where you would have all of the things that you'd regularly change together in one place. BUT, since the quality AND speed are SO good, it's great to remove anything I don't want, and put the parts together that I use all the time. The result is a fast workflow that looks better than anything I've used for Minimax. Well worth giving it a shot, at least to learn from.
I agree. The Prompt, Seed, and Model Preview Override nodes should be near each other. When testing, they are the most useful/modified nodes to see at a glance. But it's an easy fix though.
@teiji25762 the fix is I'm releasing my Minimax Seed Hunter workflow tomorrow and it's clean as fuck. It's actually already available on my patreon (for free) because I post everything early there.
Is it possible to disable the video only + png image and only keep the video with audio as the final file? My output folder is being spammed with too many of these files that I don't need.
Sounds like you need to go into your comfy settings -> VHS -> disable saving immediate files.
@foxydits Thank you!
Is anyone having luck with Temporal Upsampling De-Rope, for me it just takes up Vram
can i generate videos with or without dialogue audio using ref2va without using anything inside the audio loaders?
i tried using promps by saying something like <d> [English] blah </d> for specific words from the character, but they either utter a bit of what i want them to say then continue rambling nonsense, or its nothing but exclusive nonsense.
on the other hand, when i try to force no dialogue in the prompts (example: (no dialogue at all during the entire video:2.0)), the character spits out not but nonsensical gibberish anyway.
Theres gotta be a way around all this, right?
i dont wanna have to find another workflow. this one is amazing so far.
thanks in advance.
Glad to see this updated. Although it is buggy, the sound doesn't match the lipsync because everything seems sped up. Did t2v and img2vid. I did not use cutom audio.
How come you switched to er_sde and beta?
I tried with all the optional functions enabled and last time all were enabled except Spectrum and VFI Frame
Prompt is: Live-action, ultra-realistic cinematic footage, photorealistic skin and classroom detail, natural fluorescent lighting mixed with soft window light. Duration: 5 seconds A fixed shot nude woman on a desk in a classroom in contorsionist pose, she performs oral sex on herself. Timeline: 0.0 – 3.0s: She licks her vulva repeatedly. She looks around the class with a knowing smirk. 3.0 – 5.0s:She starts speaking in a sweet but filthy tone: “So, who else wants to taste?” Sound: Classroom ambience, quiet student murmurs and suppressed laughter.
+5 derope steps
In the example image, this means the original 10 steps become 15 steps, and the denoise is reduced to 0.4. Is my understanding correct?I don't get why Civitai won't let me add this as a resource that I used, but here's a post of a video that this workflow was VERY helpful for. https://civitai.red/posts/30547876
The outputs are so high quality.
You can add posts to the workflow right above the gallery images on the workflow page
@ODSTgolfzulu No, you can only make a new post that way. I can't add posts I already made.
Looks great!
@foxydits I still don't know what you're doing, but you fixed up the clarity of the chains when I couldn't do that in any other way.
@Jellai ComfyUI is voodoo for sure. I will say this: having 20 years in programming barely helped. The logic puzzles that comfyUI creates are entirely unique.
hi, for some reason i can't get it working, it keep showing "error:can't access property "output", res is undefined" every time i hit run button.
That error's giving me the impression that one of the nodes the workflow uses can't be reached. I can't really comment too much further since this workflow isn't maintained anymore, replaced by my new Seed Hunter workflow: https://civitai.red/models/2881362/minimax-seed-hunter-workflow-optimized-fast-latent-upscaler-speedups
thanks for the reply, i will check it out.
I truly don't understand why the 'sparse attention' thing literally makes zero difference on any videos I try to create, when everyone else is saying it makes such a huge difference for them. Everything is updated, all nodes are install, I have a pretty capable blackwell card, but still nothing. Basic ck attention with no additional speed up options is usually quicker for me. I dont get it and it makes me sad, lol. :(
This is the only workflow that gets me good H3 results. Well done! But when I try to do video continuation using a reference video and the last frame as an image anchor, it's still not seamless. Is there a way to add better shot/video continuity if I want to have a shot last a lot longer than 15 seconds, using the video/audio of the earlier video? I know for separate shots, a video editor is the way to go. But for continuing a single shot, is there a way to do it?
My newest workflow supplants this one. I'm glad you get results but my Seed Hunter workflow is the one I intend to maintain going forward: https://civitai.red/models/2881362/minimax-seed-hunter-workflow-optimized-fast-latent-upscaler-speedups
As for seamless continuations, Minimax's built in method for is garbage and always will translate/distort to the starting video. The best way to achieve it is with the Add Guide node. It's not implemented yet in either workflow but I plan to today (in the new workflow, of course).
Can´t run it with 0.4 MP, 6 Stepsin a 4090 with 24VRAM and 128Gb RAM.
Ref2Va int4
vae
32B Qwen3vl
LightX2V_flv_v1
The process gets stuck at the beginning and show nothing.
Same models used in other workflows return a good result.
I´m missing something?
Solid workflow. The little time I had to play with it, it works fantastic. Very well laid out.
Just wait until you try the current one (Minimax Seed Hunter). This workflow has been shelved and replaced with the aforementioned Seed Hunter version. I'm glad you enjoyed this one though!
@foxydits Will it make my 5090 scream like a banshee? 😂
@mattoneplus12818 Depends on how much you wanna upscale your latent by ;)

