Simply Advanced - MiniMax H3 Workflow by altoiddealer
This workflow was built from the ground up by myself working daily over the course of 2-3 weeks, aiming to be well structured, expandable, and easy to use without sacrificing any functionality.
New in v1.5:
Seamless Continuation has been replaced with Seamless Extension
Slices off X frames from your video input(s) to set Video/Audio guide(s).
MiniMax H3 generates Video/Audio which begins and/or ends with those sliced inputs.
The workflow stitches the result with the original video input(s) using configurable overlap blending strategies. Optional: enter the subgraph to enable saving the raw generation result.
Seamless Extension modes:
Seamless Continuation (generate from input video as start point)
Seamless Convergence (generate into input video as end point)
Can be used in combination (Generate between two input videos)
Main features:
Supports up to: 6 Images, 2 Videos, 3 Audio.
ANY can be set as reference, guide, or both. This is much more useful than the average user understands - I have not yet seen this support in any other workflow.
Latent Upscale and/or 2-Stage Generation
Intuitive settings/LoRA management for each stage.
1-4 steps LLM Prompt Enhancement
Subgraph contains a dashboard to generate / edit LLM's responses before they pass into the next stage.
Includes my personal 4-Step Ref2VA pipeline.
Step 1 Prompt Enhancement/Expansion: (Optional Step). Allows the LLM to interpret the user's prompt more verbosely and in relation to the reference materials. Catching/editing discrepancies here will yield much better results in next steps.
Step 2: Reference Analysis: Creates a detailed mapping of the references in relation to the Expanded Prompt. The resulting Reference Map is included as part of the context in the next steps.
Step 3: Creative Director: Considers Expanded Prompt + Reference Analysis to create a Creative Blueprint.
Step 4: Prompt Compiler: Conforms the Reference Map + Creative Blueprint into the H3 prompt structure
Includes my personal 1-Step FL2VA Prompt Enhance
The system prompt begins with base instructions, adding task specific instructions (T2VA / I2VA / FL2VA / L2VA)
The mode is detected automatically based on the state of Image 1 / Image 2.
You, the user, can easily add your own LLM instructions into this system as selectable options. Take a shot at making your own multi-stage pipeline!
Seamless video extension option - continuation idea taken from Seed Hunter workflow - later expanded to include convergence/combination usage.
Force Audio option - idea taken from Seed Hunter workflow
Easily expandable framework
Easily add/manage speedups, LoRAs, etc
Is the poster child for an amazing new node, Pass or None. Workflow developers, start using this!!

ComfyUI Requirements:
Be updated.
MUST have comfyui-frontend version < v1.50.5 (Fixed broken Custom Combo nodes in subgraphs)
Required Custom Nodes:
comfyui_essential-er (My repo). Was optional, now Required for v1.5+ for its superior video/audio merging node
Optional Custom Nodes:
Load Video (Upload) from ComfyUI-VideoHelperSuite Feel free to replace with another Load Video node.
Datetime String from comfyui-various - Useful for naming outputs. NOT REQUIRED
ComfyUI-Spectrum-MiniMax-H3 - Speedup option. NOT REQUIRED
Description
v1.5.4: Fixed force_audio not actually working. Fixed optional match image input being disconnected in <Picture 2>.
v1.5.3: Reverted the chain sampling implementation from 1.5.2. Currently too experimental.
v1.5.2:
D'ya know the expression, "you have to break a few eggs to make an omelet"? This fixes an issue where the "Second Pass" toggle had an inverted effect
Disabled = Yes / Enabled = No. FIXEDAdded improved chain sampling logic, so you can optionally return leftover noise from a not fully denoised First Pass.
v1.5.1: Fix error when not using video inputs. Sorry to any who snatched v1.5 right away and experienced this.
v1.5:
ComfyUI-essential-er (my repo) is now required.
Reason: Extremely useful video/audio merging node which KJNodes has failed to merge from both a PR I had opened, as well as another user's still open PR.
In terms of this workflow, it turns this (v1.4) into this (v1.5).
New Feature:
"Seamless Continuation" has been replaced with "Seamless Extension"
Seamless Extension modes:
Seamless Continuation (generate from your video onwards)
Seamless Convergence (generate into your source video)
Can be used in combination
FAQ
Comments (4)
Seamless Extension: My 1 example in showcase was done lazily with 3 steps + Taomate LoRA. I used merely 22 frames of video/audio on each side. The workflow automatically slices it off from the input videos, setting them as Guides. The target output video was 6 seconds so it generated ~4 seconds of "new" video/audio, which this workflow stitches back onto the original videos resulting in one long video.
This looks awesome! Do the references work for i2v or do they only work with ref2va?
I've been spending most of my time developing the workflow than actually using it. However, I can say from my limited testing that References are factored when using the FL2VA model - as I just used it yesterday including an audio reference and it was effective. I've seen in comments on Reddit, Discord, etc - that the FL2VA model in many cases actually works better than the Ref2VA model for reference oriented generations. I'm speaking from practically zero personal experience though.
What I did test a few times was using the Ref2VA model / FL2VA model -- in this workflow setting Image1 and/or Image2 as Guides only -- it behaves just like using the FL2VA workflow (even though the node applying the conditioning is the "Ref2VA" one opposed to the dedicated "FL2VA" (the one with the first_image and last_image inputs). I did not deep dive into the code but I imagine that the FL2VA node simply just sets Guides - which this workflow manually sets based on the little panel next to the image inputs.
@altoiddealer That's been my experience. FL2VA or hybrid models work better than dedicated ref2VA models.



