CivArchive
    Preview 144366706
    Preview 144366707

    MiniMax H3 LongForge v4.0 — FL2VA & REF2VA

    My Telegram Channels

    AI Bob Public
    AI DEV LAB

    Support the project

    If LongForge helps you, you can support future development through the Tip button on my Civitai profile or with an optional donation:

    BNB Smart Chain

    0xdCf9846A1B0F2fe32E19dC9CFdDC5194CC13D108

    TON

    UQCeukjBulZw9gni1jv8NLRgtmgHDQ6mUiS2qf3716XobIxA

    TRC20

    THLU3yu8VpoPqP4bMLCp86FRhoCweJ2Qz4

    Please send only USDT through the TRON (TRC20), TON, BNB SmartChain network. Thank you!

    One editable ComfyUI workflow for generating, continuing, enhancing, reviewing, recovering and exporting long MiniMax H3 videos with audio.

    LongForge transfers native video and audio latent context between scenes, helping preserve motion, character appearance, camera movement, scene geometry and ambience.

    Version 4.0 introduces optional SOFT audio latent continuation: most of the copied audio overlap remains protected, while a short transition region lets H3 gradually rework its ending. Film export uses that reworked region when assembling the continuous soundtrack.

    The release also reorganizes advanced audio controls, updates settings restoration and includes FFmpeg executable validation for previews and export.

    The local Prompt Assistant, automatic overlap selection, Adaptive Native Cache, HQ Finish, clean refiner branch, persistent projects, recovery metadata and long-audio references remain available in the same workflow.

    Important: Turbo LoRA and Adaptive Native Cache are separate acceleration methods. Keep Native Cache bypassed when using Turbo or another short-step acceleration setup.

    What’s new since v3.7

    SOFT audio latent continuation

    The new SOFT mode gives H3 a short transition region at the end of the copied audio overlap.

    Earlier audio context remains protected exactly. Inside the transition region, a gradual cosine mask allows H3 to rework the saved audio latent as it generates the next scene.

    Choose between:

    Soft transition widthDuration4 audio latent tokens100 ms8 audio latent tokens200 ms

    These values count temporal audio latent tokens, not sampler steps.

    The transition starts from the real saved audio latent and happens during native H3 generation. The copied video overlap remains hard-protected.

    SOFT provides another way to handle audio continuation. Results depend on the checkpoint, prompts and generation settings; compare it with HARD for your film.

    Audio continuity — OFF / HARD / SOFT

    Audio continuity replaces the previous Preserve audio overlap switch.

    ModeBehaviorOFFLeaves audio open to generation while preserving the video overlap.HARDCopies and protects the entire audio overlap exactly, preserving the previous enabled-switch behavior.SOFTProtects the earlier audio context and gradually allows H3 to rework a short tail inside the overlap.

    HARD remains the default in the supplied workflow. Select SOFT explicitly when you want to try the new transition.

    Legacy settings retain their original meaning: an enabled audio-preservation switch maps to HARD, and a disabled switch maps to OFF.

    Film export supports the reworked audio tail

    When exporting SOFT continuations, LongForge uses the next scene’s reworked soft tail in the assembled soundtrack, followed by its newly generated audio.

    The complete assembled audio latent is decoded once as a continuous timeline.

    This preserves:

    • the film’s duration;

    • the video cut positions;

    • the original saved parent takes;

    • exact protection of the hard audio core.

    Changing Audio continuity affects newly generated continuations. Existing saved scenes keep their original audio. Regenerate from the affected scene to change its transition.

    Use the v4.0 exporter for SOFT takes.

    Cleaner audio controls

    Less frequently used controls now sit behind two buttons:

    ButtonControlsAdvanced audioSOFT transition width and Extra audio history.Advanced audio exportRepair audio clicks and Diagnostic WAVs.

    The SOFT width appears only when SOFT is selected. Extra audio history remains a separate setting.

    Repair audio clicks stays enabled by default, and Diagnostic WAVs stay OFF.

    These changes use the existing nodes and connections.

    Export and settings restoration

    Preview and export check that the selected FFmpeg executable is usable and provides the required H.264 and AAC encoders. An incompatible command found on PATH can be skipped in favor of the imageio-ffmpeg fallback.

    Settings restoration also accounts for the previous and current video-control layouts, including the existing automatic continuity field.

    Updated public workflow and guides

    The supplied workflow contains a neutral boat example, an empty Reference Library and placeholders for model selections.

    Both LoRA loaders and both FL2VA image loaders start bypassed. KJ attention patches and Native Cache / BALANCED are active, HQ Finish is OFF, Audio continuity is HARD, and Video metadata is NONE.

    The English/Russian guide covers SOFT continuation, its export behavior, advanced audio controls, compatibility and the supplied v4.0 configuration.

    Main features

    • Unified FL2VA and REF2VA workflow.

    • Persistent multi-scene projects.

    • Native video and audio latent continuation.

    • OFF / HARD / SOFT audio continuity.

    • Optional 100 / 200 ms soft audio transition.

    • Manual and automatic overlap selection.

    • Scene Studio with searchable prompts, statuses and previews.

    • Optional local prompt drafting with Gemma 4.

    • Drag-and-drop reordering of pending scenes.

    • Built-in image, video and audio Reference Library.

    • Automatic long-audio slicing for REF2VA.

    • Scene regeneration and controlled variations.

    • Optional recovery metadata inside preview and export files.

    • Integrated native latent upscaling and refinement.

    • Separate clean model branch for refinement.

    • Adaptive Native Cache with manual and AUTO profiles.

    • Local audio-click repair, optional video-seam correction and join diagnostics.

    • Complete film export with a continuous generated soundtrack.

    • Per-project locking and atomic updates on Windows.

    Generation modes

    FL2VA

    Supports text-to-video, image-to-video, optional first-frame guidance, optional last-frame guidance and combined first/last-frame generation.

    Leave both image loaders bypassed for text-only generation. First frame anchors the film’s beginning. Last frame applies to selects either the last scene currently listed or every scene.

    REF2VA

    Supports up to:

    • 9 image references;

    • 3 video references with optional paired audio;

    • 3 standalone audio references.

    Switch the pipeline inside the same workflow and select the matching MiniMax H3 diffusion model.

    FIRST SCENE presents ordinary references only at the beginning; later scenes continue from saved latent context.

    EVERY SCENE presents them again and can pull generation back toward the original pose, composition or sound.

    Follow-scenes audio advances independently of this policy.

    Optional local Prompt Assistant

    Write a short scene idea directly in Project & Scenes, then press Generate draft to create an English H3 prompt with a separate local Gemma 4 text model.

    The assistant uses:

    • the selected FL2VA / REF2VA pipeline;

    • the configured scene duration;

    • active first/last-frame inputs;

    • available reference labels;

    • the previous scene’s prompt for continuation.

    Review or edit the result, then choose Apply to scene. Closing the draft window leaves your scene unchanged.

    Applying a draft to an already-saved scene changes its editor text. Regenerate that scene to update the video.

    The bundled base and reference prompt-writing guides provide the required structure. Validation checks required fields, their order, unavailable reference labels and some incorrect reference-use claims.

    The assistant receives text and reference labels, not the actual images, video frames or audio. Describe important source features in your idea.

    In REF2VA, a picture is an appearance/style reference by default. Explicitly request a first, last or intermediate keyframe when that is the intended use.

    Gemma is optional and separate from the H3 text encoder. Normal video generation and export do not require it.

    Automatic continuity modes

    Video & Continuity offers three overlap-selection modes:

    ModeBehaviorMANUALUses your selected overlap exactly. The supplied workflow starts at 22 frames.AUTO BALANCEDEvaluates the previous native AV tail and normally chooses 5 / 22 / 39 frames.AUTO SAFEUses the more conservative 22 / 39 / 56-frame pool when sufficient context and scene length are available.

    AUTO examines changes in the saved video latent and, when audio continuity is enabled, the audio latent. It selects a valid overlap for the next scene without rewriting the previous saved take.

    The actual overlap and selection reason are saved with each scene and shown in its diagnostics.

    In AUTO modes, the overlap field displays that automatic selection instead of acting as a manual override.

    Automatic selection helps choose context length; it does not guarantee an invisible or inaudible join.

    Adaptive Native Cache

    Adaptive Native Cache reuses selected stable model outputs while keeping the sampler schedule intact. It is intended for native H3 generation, normally 20 or 25 steps with CFG 1.0.

    ProfilePurposeAUTOScene-specific calibration, fresh-output checks and automatic fallback.QUALITYThe most conservative manual reuse profile.VOICEStricter audio protection for speech and vocals.ACTIONCautious reuse for motion-heavy scenes.BALANCEDGeneral starting point and supplied workflow default.FASTMore aggressive reuse.PREVIEWFaster drafts and prompt testing.DRAFT10Targets roughly ten reused steps in a 20-step run when checks allow.CUSTOMManual thresholds, budgets, warmup, reuse window and consecutive-reuse limits.

    Custom controls appear only when CUSTOM is selected.

    AUTO profile

    AUTO calibrates on the current scene and selects an existing manual profile according to video/audio behavior.

    Before allowing its first reuse, and periodically afterward, it computes a fresh model output and checks the proposed cached result against it.

    Failed checks move the controller toward a more conservative profile. Repeated failures can disable reuse for the rest of that scene. Verification steps keep the fresh result.

    Compatibility and comparisons

    Video and audio are evaluated separately; audio can veto reuse. Warmup, the final step and protected cache-window edges use full computation.

    Reuse is disabled for fewer than 16 steps, CFG other than 1.0 and detected incompatible configurations. Compatibility guards also cover recognizable acceleration LoRAs and stochastic sampler variants.

    Do not stack another output-cache system on the same model branch. A renamed acceleration checkpoint may require manual bypass.

    Cache operates during generation. It does not replace latent continuation or process HQ refinement, preview decoding or film export.

    Reports show actual full/reused model calls. Estimated time savings are not a measured speedup guarantee.

    Bypass cache for Turbo, short-step acceleration or an exact native comparison. Compare a chosen profile against bypass with the same seed before relying on it for final output.

    Integrated HQ Finish

    The optional HQ Finish node processes each completed scene before preview and export.

    ModeProcessingOFFNative H3 output.UPSCALELearned MiniMax H3 3D latent upscaling.UPSCALE + REFINELatent upscaling followed by low-strength native H3 video refinement.

    The integration is built into LongForge. Download a compatible third-party 3D-conv upscaler checkpoint separately; no separate upscaler node pack is required.

    Clean refiner branch

    The supplied workflow has two model paths:

    • Generation: shared attention patches → Turbo LoRA → Extra LoRA → optional Adaptive Native Cache.

    • Refinement: shared attention-patched H3 model before both LoRAs and cache.

    This keeps generation LoRAs and cache out of the refinement pass while retaining the shared attention patches.

    Recommended starting settings:

    • Refine steps: 5

    • Refine denoise: 0.25

    • Sampler: Euler

    • Scheduler: Simple

    • CFG: 1.0

    Five passes at 0.25 denoise sample the final part of a native 20-step trajectory. Other useful pairs are 4 / 0.20, 6 / 0.30 and 7 / 0.35.

    Native continuation and HQ export

    LongForge preserves the native video/audio latent for continuation and stores enhanced HQ video for previews and export.

    Audio is not upscaled and remains protected during refinement. The previous HQ overlap is carried into the next HQ scene.

    Keep the same HQ mode, checkpoint, precision and output geometry throughout the film. Regenerate from Scene 1 when changing these settings.

    For example, a film generated at 736 × 416 with a 2.0× HQ scale produces 1472 × 832 video while continuation still uses the original native latent.

    HQ increases processing time and storage requirements.

    Scene Studio

    Project & Scenes provides editable scene cards, search, generation statuses, saved previews, pending-scene reordering, regeneration, variation and project cleanup controls.

    Pending cards can be dragged into a new order while search is empty. Saved scenes depend on the latent context before them; changing an already-generated part requires rebuilding its continuation.

    Open project restores saved prompts and video, sampling and HQ settings. Model and LoRA selections remain in the workflow JSON.

    Save that JSON before closing the browser to retain pending editor changes and loader selections.

    Use a different project name for a separate film.

    New film resets the active history for the current name while retaining editor text and old takes. Clean unused removes takes outside the active chain; Delete film removes the entire project.

    Audio continuity

    Starting settings in the supplied workflow:

    SettingValueContinuity modeMANUALOverlap22 framesAudio continuityHARDSOFT width4 latent tokens / 100 ms, used only in SOFT modeExtra audio history0

    Select AUTO BALANCED or AUTO SAFE when you want LongForge to choose overlap from the preceding scene.

    A 5-frame overlap provides less context and a higher risk of motion or audio resets. For SOFT comparisons, start with 22 frames.

    Trying SOFT

    1. Select Audio continuity → SOFT.

    2. Open Advanced audio.

    3. Start with 4 audio latent tokens / 100 ms.

    4. Generate or regenerate the continuation.

    5. Choose EXPORT FILM and listen across the join.

    Try 8 tokens / 200 ms when you want a longer transition region.

    SOFT must leave a nonempty hard-protected audio core. A very short 5-frame overlap with an 8-token tail can fail this requirement, depending on its timeline position. LongForge rejects that configuration before sampling.

    Normal scene previews contain only the newly added frames and omit the overlap. Evaluate the SOFT transition in the exported film or its diagnostic WAVs.

    For a direct comparison, use separate projects with the same prompts, seeds, model and geometry. Bypass Native Cache and disable Repair audio clicks while comparing HARD, SOFT 4 and SOFT 8.

    Review ambience, sound levels, speech continuity and lip sync. Keep the mode that works best for the chosen film.

    Extra audio history

    Extra audio history supplies additional context immediately before the protected overlap. It is independent of the SOFT transition width and is not repeated in the film.

    Start at 0. Custom values use multiples of 3 frames, and the history plus actual overlap must fit inside the source scene.

    If active REF2VA audio references conflict with this additional history, LongForge disables only Extra audio history for that scene and reports it. Audio overlap protection remains enabled.

    Long-audio references

    In REF2VA, choose Audio use → Follow scenes · REF2VA in the audio card.

    LongForge advances through the source according to the saved film timeline, avoiding a duplicate reference for audio already covered by the protected overlap.

    This guides H3’s own audio synthesis. It does not replace the final soundtrack or guarantee exact words, voices, rhythm or timing.

    Actual source intervals appear in the generation report. If the selected source ends early, the remaining reference window is padded with silence.

    Seam repair and diagnostic WAVs

    Open Advanced audio export to access Repair audio clicks and Diagnostic WAVs · export only.

    Repair audio clicks applies a bounded correction around detected isolated clicks near joins. A broad dropout, missing sound or sustained reset already generated by H3 may require scene regeneration.

    Diagnostics examine both audio channels and distinguish short transients from sustained changes. Export reports identify joins requiring review as before Scene N.

    For an audible join, enable Diagnostic WAVs · export only, select EXPORT FILM and run again. No new diffusion generation is required.

    The export folder receives:

    • the MP4;

    • .before_repair.wav — continuous decoded H3 audio before seam repair;

    • .before_aac.wav — audio after seam repair and before AAC encoding.

    The WAVs retain float32 samples without lossy compression. Both reflect the Fade film edges setting.

    Compare files with the same export basename around the reported join. Enabling diagnostic WAVs alone does not repair the sound.

    Repair video seams is optional and OFF by default.

    Fade film edges applies short ramps only at the beginning and end of the complete film.

    Recovery metadata

    Set Video metadata → RECOVERY to embed prompts and workflow recovery information in newly created previews and exported MP4s.

    This can include model/LoRA filenames and reference-media selections.

    Drag a recovery MP4 into ComfyUI to reopen its workflow. A scene preview restores that scene’s graph; a full-film export uses the last scene’s model setup with the active film prompts.

    It does not automatically switch a different LoRA setup for every card.

    Recovery does not contain model weights, LoRA files, original reference media or scene latents. Keep the project folder to continue generating after a restart:

    ComfyUI/output/longforge_native/<project>/

    The public workflow starts with Video metadata → NONE. Select RECOVERY when you want embedded restoration data.

    NONE affects newly written videos. It does not remove metadata from existing MP4s or delete project data. ComfyUI’s --disable-metadata also disables embedded recovery.

    If recovery export fails because of a JSON nan value, select NONE. This does not change generation itself.

    Project handling on Windows

    LongForge uses per-project locking, atomic JSON updates and protection against simultaneous generation and cleanup.

    Successfully generated scenes are saved immediately, including during ALL PENDING.

    Clean unused and Delete film require confirmation. Avoid editing or generating the same project simultaneously in several browser tabs.

    Existing compatible project/take folders can be reused. Install the supplied v4.0 workflow to access the current controls and clean refiner connection.

    Previously saved takes retain their original transitions. Regenerate the affected continuation when switching it to SOFT.

    Workflow structure and defaults

    The graph uses visible native model, CLIP, VAE and LoRA loaders; KJNodes attention patches; and LongForge project, continuity, reference, cache, HQ, prompt-assistant and export nodes.

    ComponentSupplied v4.0 settingPipelineFL2VAResolution1344 × 768Scene length158 frames at 24 FPSContinuity modeMANUALOverlap22 framesAudio continuityHARDSOFT width4 latent tokens / 100 ms; active only in SOFT modeExtra audio history0Sampling20 steps, res_multistep, simple, CFG 1, denoise 1Turbo LoRA / Extra LoRABypassed; no files selectedFirst / Last frame loadersBypassed; no files selectedKJ attention patchesEnabledAdaptive Native CacheEnabled, BALANCEDHQ FinishOFFPrompt AssistantOptional; text model not connectedReference LibraryEmptyVideo metadataNONERepair audio clicksONRepair video seamsOFFFade film edgesONDiagnostic WAVsOFF

    Sampling remains editable. Euler/simple is a straightforward starting point for native sampling and cache comparisons; the supplied JSON retains res_multistep as its selected sampler.

    Width and height must be multiples of 32. Scene length and overlap follow the H3 grid 5 + 17 × k frames, with overlap shorter than the next scene.

    There is no fixed 1 MP resolution cap; practical limits depend on the model and hardware.

    Scene duration includes overlap. Three 243-frame scenes with 22-frame overlap produce:

    243 + 221 + 221 = 685 unique frames

    That is approximately 28.54 seconds at 24 FPS. Selecting SOFT does not change this duration.

    Download

    Download ComfyUI-H3-LongForge-v4.0.zip.

    The archive contains:

    • ComfyUI-H3-LongForge-NodePack/ — custom nodes, web interface, bundled prompt-writing guides and installation README.

    • H3_LongForge_FL2VA_REF2VA_v4.0.json — the complete FL2VA / REF2VA workflow.

    • PROMPT_GUIDE_RU_EN_v4.0.md — the complete English/Russian user and prompt guide.

    • START_HERE_v4.0_RU.md — Russian quick-start instructions.

    Models, LoRAs, VAEs, Gemma weights and upscaler checkpoints are not included.

    LongForge v4.0 requires:

    • a compatible ComfyUI build with native MiniMax H3 support and the V3 node API;

    • a matching H3 FL2VA or REF2VA diffusion model;

    • a compatible H3 text encoder;

    • H3 video and audio VAEs;

    • a usable FFmpeg executable or the imageio-ffmpeg fallback.

    The supplied graph also uses ComfyUI-KJNodes and compatible SageAttention for its enabled attention patches.

    These patches are optional to LongForge’s core pipeline. Install their dependencies to use the graph as supplied, or remove the patch nodes and reconnect the model paths.

    Optional features require matching LoRAs, a compatible 3D latent-upscaler checkpoint, or a separate Gemma 4 text model for Prompt Assistant.

    MiniMax H3 models

    Comfy-Org / MiniMax-H3

    Alternative H3 text encoders

    INT8 ConvRot · NVFP4

    Choose a format compatible with your ComfyUI build and hardware. These encoders serve H3 conditioning; Prompt Assistant uses its own text-generation model.

    Optional Prompt Assistant model

    Comfy-Org / Gemma 4 text encoders

    The bundled guide lists gemma4_12b_int8_convrot.safetensors and gemma4_e4b_it_fp8_scaled.safetensors as options for a ComfyUI build supporting them.

    Attention nodes

    ComfyUI-KJNodes

    Optional latent upscaler

    LBH-123-AI / MiniMax H3 Latent Upscaler

    Use a compatible minimax_h3_latent_upscaler_3d_conv_v1_*.safetensors checkpoint. This integration supports the 3D version.

    Reference-media loading and upscaler integration are built into LongForge; neither requires a separate node pack.

    Installation — Windows Portable

    1. Close ComfyUI completely.

    2. Remove the previous ComfyUI-H3-LongForge-NodePack folder. Keep your projects under ComfyUI/output/longforge_native/.

    3. Extract the node-pack folder from the v4.0 archive so this file exists: ComfyUI/custom_nodes/ComfyUI-H3-LongForge-NodePack/__init__.py.

    4. Keep the workflow and guides from the archive separately. Do not merge node-pack versions or leave duplicate LongForge installations in custom_nodes.

    5. Install/update ComfyUI-KJNodes and compatible SageAttention for the supplied attention nodes.

    6. If your installation needs the FFmpeg fallback, run the following from ComfyUI_windows_portable:

      python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-H3-LongForge-NodePack\requirements.txt
    7. Restart ComfyUI and refresh the browser with Ctrl+F5.

    8. Open H3_LongForge_FL2VA_REF2VA_v4.0.json.

    9. Select the H3 model, H3 text encoder, video VAE and audio VAE. Match Pipeline to the model.

    10. Select optional LoRA or image files before enabling their loaders with Ctrl+B. Bypass Native Cache when using Turbo.

    Optional Prompt Assistant setup

    1. Use a ComfyUI build with native Gemma 4 and text-generation support.

    2. Download a compatible Gemma encoder into ComfyUI/models/text_encoders/ and restart ComfyUI.

    3. Add a native Load CLIP node in the Prompt Assistant area.

    4. Select the Gemma file and set type = stable_diffusion.

    5. Connect its CLIP output to Prompt Assistant → text_model.

    6. Write a short idea in a scene card, press Generate draft, review the result and choose Apply to scene.

    For several connected scenes, apply Scene 1’s draft before drafting Scene 2, then apply Scene 2 before drafting Scene 3.

    Each request receives the preceding card’s text. Review the actual generated scene separately when deciding how its continuation should develop.

    Optional HQ Finish setup

    1. Download a compatible 3D-conv latent-upscaler checkpoint.

    2. Place it in ComfyUI/models/latent_upscale_models/.

    3. Restart ComfyUI and select the checkpoint in HQ Finish.

    4. Choose UPSCALE or UPSCALE + REFINE.

    5. For refinement, retain the supplied clean refiner connection before both LoRAs and cache.

    Generate your film

    Give the project a unique name. Write one prompt per scene and use + Scene to add the next continuation.

    A simple manual prompt structure:

    VIDEO: Describe the subject, action, setting and camera.
    SOUND: Describe ambience, effects and dialogue.
    MUSIC: No music.

    These headings are optional. Full native H3 prompts, including the longer structure produced by Prompt Assistant, can also be used.

    Continuation prompts should describe the next stage of the same action while preserving the intended subject, setting, camera and sound.

    Do not put several repeated VIDEO sections in one card and expect separate scenes.

    REF2VA reference markers

    Use <Picture 1>, <Video 1> and <Audio 1> for available references.

    Optional aliases can make mixed inputs easier to address:

    hero=image1
    motion=video1
    score=audio1
    voice=video1_audio

    Then use {hero}, {motion}, {score} or {voice} in the prompt.

    Define only sources you actually loaded. With FIRST SCENE, later prompts should continue the generated scene without addressing static references that are no longer being supplied.

    Main actions

    Action Result

    GENERATE NEXT + ONE SCENE -> Generates the next pending scene.

    GENERATE NEXT + ALL PENDING -> Generates all remaining scenes in order, saving each one.

    REGENERATE SELECTED -> Rebuilds from the selected saved scene with its saved seed.

    REGENERATE LAST -> Rebuilds the last saved scene with its saved seed.

    VARY SELECTED -> Rebuilds the selected scene with its saved seed plus one.

    PREVIEW SELECTED -> Recreates a missing preview from saved latents without diffusion.

    EXPORT FILM -> Assembles the active saved scenes into one MP4 with generated audio.

    When regenerating, ONE SCENE replaces only the chosen scene and leaves its continuation pending. ALL PENDING also rebuilds the following scenes.

    If all cards are already saved, GENERATE NEXT regenerates the selected scene. Add a new card first when you want to extend the film. Export uses saved scenes and does not generate pending cards.

    When regenerating, ONE SCENE replaces only the chosen scene and leaves its continuation pending. ALL PENDING also rebuilds the following scenes.

    If all cards are already saved, GENERATE NEXT regenerates the selected scene. Add a new card first when you want to extend the film.

    Export uses saved scenes and does not generate pending cards.

    For SOFT continuations, review the complete exported join: the isolated scene preview omits the overlap containing the soft transition.

    Optional final post-processing

    For separate enhancement or frame interpolation after export:

    DLSS 5 Visual Enhancer · Downloads

    This is a separate application with its own requirements and is not required by LongForge.

    ===========================================================

    MiniMax H3 LongForge v3.7 — FL2VA & REF2VA

    My Telegram Channels

    AI Bob Public
    AI DEV LAB

    Support the project

    If LongForge helps you, you can support future development through the Tip button on my Civitai profile or with an optional donation:


    BNB Smart Chain
    0xdCf9846A1B0F2fe32E19dC9CFdDC5194CC13D108
    TON
    UQCeukjBulZw9gni1jv8NLRgtmgHDQ6mUiS2qf3716XobIxA

    TRC20
    THLU3yu8VpoPqP4bMLCp86FRhoCweJ2Qz4

    Please send only USDT through the TRON (TRC20), TON, BNB SmartChain network. Thank you!

    One editable ComfyUI workflow for generating, continuing, enhancing, reviewing, recovering and exporting long MiniMax H3 videos with audio.

    LongForge transfers protected video and audio latent context between scenes, helping preserve motion, character appearance, camera movement, scene geometry and ambience.

    Version 3.7 adds an optional local Prompt Assistant, automatic continuity modes, an AUTO profile for Adaptive Native Cache, and more detailed audio and scene-join diagnostics. HQ Finish, the clean refiner branch, persistent projects, recovery metadata and long-audio references remain part of the same workflow.

    Important: Turbo LoRA and Adaptive Native Cache are separate acceleration methods. Keep Native Cache bypassed when using Turbo or another short-step acceleration setup.

    What’s new since v3.5

    Optional local Prompt Assistant

    Write a short scene idea directly in Project & Scenes, then press Generate draft to create an English H3 prompt with a separate local Gemma 4 text model.

    The assistant uses:

    • the selected FL2VA / REF2VA pipeline;

    • the configured scene duration;

    • active first/last-frame inputs;

    • available reference labels;

    • the previous scene’s prompt for continuation.

    Review or edit the result, then choose Apply to scene. Closing the draft window leaves your scene unchanged. Applying a draft to an already-saved scene changes its editor text; regenerate that scene to update the video.

    The bundled base and reference prompt-writing guides provide the required structure. Validation checks required fields, their order, unavailable reference labels and some incorrect reference-use claims.

    The assistant receives text and reference labels, not the actual images, video frames or audio. Describe important source features in your idea. In REF2VA, a picture is an appearance/style reference by default; explicitly request a first, last or intermediate keyframe when that is the intended use.

    Gemma is optional and separate from the H3 text encoder. Normal video generation and export do not require it.

    Automatic continuity modes

    Video & Continuity now offers three modes:

    ModeBehaviorMANUALUses your selected overlap exactly. The supplied workflow starts at 22 frames.AUTO BALANCEDEvaluates the previous native AV tail and normally chooses 5 / 22 / 39 frames.AUTO SAFEUses the more conservative 22 / 39 / 56-frame pool when enough context and scene length are available.

    AUTO examines changes in the saved video latent and, when audio preservation is enabled, the audio latent. It selects a valid overlap for the next scene without rewriting the previous scene. Very short windows can require smaller valid values.

    The actual overlap and selection reason are saved with each scene and shown in its diagnostics. Manual controls remain available. Automatic selection helps choose context length; it does not guarantee an invisible or inaudible join.

    Adaptive Native Cache — AUTO profile

    The new AUTO profile calibrates on the current scene and selects an existing manual profile according to video/audio behavior.

    Before allowing its first reuse, and periodically afterward, AUTO computes a fresh model output and checks the proposed cached result against it. Failed checks move the controller toward a more conservative profile; repeated failures can disable reuse for the rest of that scene. Verification steps always keep the fresh result.

    The existing QUALITY / VOICE / ACTION / BALANCED / FAST / PREVIEW / DRAFT10 / CUSTOM profiles remain available. The supplied workflow uses BALANCED; AUTO is an optional selection.

    Additional compatibility guards detect recognizable acceleration/Turbo LoRAs in the active generation branch and disable reuse for stochastic samplers such as ancestral, SDE, DDPM and LCM variants. A renamed acceleration checkpoint may require manual bypass.

    Audio export and join diagnostics

    Audio work focuses on identifying the source of a problem and keeping repairs local:

    • Repair audio clicks applies a bounded correction around detected isolated clicks instead of smoothing a large part of the soundtrack.

    • Diagnostics examine both audio channels and distinguish a short transient from a sustained level reset.

    • Export reports identify joins that need review as before Scene N.

    • Diagnostic WAVs · export only saves the continuous decoded audio before seam repair and the final PCM before AAC encoding.

    • FFmpeg export uses the film’s timeline duration to avoid premature termination that could cut off a trailing AAC packet.

    Broad dropouts, missing sound and resets already generated by H3 cannot be reconstructed during export. Diagnostic WAVs help determine where a transition appeared; enabling them alone does not remove it.

    Updated public workflow and guides

    The supplied workflow contains a neutral boat example, an empty Reference Library and placeholders for model selections. Both LoRA loaders and both FL2VA image loaders start bypassed. The KJ attention patches and Native Cache / BALANCED are active, HQ Finish is OFF, and Video metadata is set to NONE.

    The English/Russian guide now covers Prompt Assistant setup, sequential scene drafting, automatic continuity, cache AUTO, audio diagnostics and the actual v3.7 starting configuration.

    Main features

    • Unified FL2VA and REF2VA workflow.

    • Persistent multi-scene projects.

    • Protected video and audio latent continuation.

    • Manual and automatic overlap selection.

    • Scene Studio with searchable prompts, statuses and previews.

    • Optional local prompt drafting with Gemma 4.

    • Drag-and-drop reordering of pending scenes.

    • Built-in image, video and audio Reference Library.

    • Automatic long-audio slicing for REF2VA.

    • Scene regeneration and controlled variations.

    • Optional recovery metadata inside preview and export files.

    • Integrated native latent upscaling and refinement.

    • Separate clean model branch for refinement.

    • Adaptive Native Cache with manual and AUTO profiles.

    • Local audio-click repair, optional video-seam correction and join diagnostics.

    • Complete film export with generated audio.

    • Per-project locking and atomic updates on Windows.

    Generation modes

    FL2VA

    Supports text-to-video, image-to-video, optional first-frame guidance, optional last-frame guidance and combined first/last-frame generation.

    Leave both image loaders bypassed for text-only generation. First frame anchors the film’s beginning. Last frame applies to selects either the last scene currently listed or every scene.

    REF2VA

    Supports up to:

    • 9 image references;

    • 3 video references with optional paired audio;

    • 3 standalone audio references.

    Switch the pipeline inside the same workflow and select the matching MiniMax H3 diffusion model.

    FIRST SCENE presents ordinary references only at the beginning; later scenes continue from saved latent context. EVERY SCENE presents them again and can pull the generation back toward the original pose, composition or sound. Follow-scenes audio advances independently of this policy.

    Adaptive Native Cache

    Adaptive Native Cache reuses selected stable model outputs while keeping the sampler schedule intact. It is intended for native H3 generation, normally 20 or 25 steps with CFG 1.0.

    ProfilePurposeAUTOScene-specific calibration, fresh-output checks and automatic fallback.QUALITYThe most conservative manual reuse profile.VOICEStricter audio protection for speech and vocals.ACTIONCautious reuse for motion-heavy scenes.BALANCEDGeneral starting point and supplied workflow default.FASTMore aggressive reuse.PREVIEWFaster drafts and prompt testing.DRAFT10Targets roughly ten reused steps in a 20-step run when checks allow.CUSTOMManual thresholds, budgets, warmup, reuse window and consecutive-reuse limits.

    Custom controls appear only when CUSTOM is selected.

    Video and audio are evaluated separately; audio can veto reuse. Warmup, the final step and protected cache-window edges use full computation. Reuse is disabled for fewer than 16 steps, CFG other than 1.0 and detected incompatible configurations. Do not stack another output-cache system on the same model branch.

    Cache does not change prompt conditioning, saved continuation latents, HQ refinement, preview decoding or film export. Reports show actual full/reused model calls; estimated time savings are not a measured speedup guarantee.

    Bypass cache for Turbo, short-step acceleration or an exact native comparison. Compare a chosen profile against bypass with the same seed before relying on it for final output.

    Integrated HQ Finish

    The optional HQ Finish node processes each completed scene before preview and export.

    ModeProcessingOFFNative H3 output.UPSCALELearned MiniMax H3 3D latent upscaling.UPSCALE + REFINELatent upscaling followed by low-strength native H3 video refinement.

    The integration is built into LongForge. Download a compatible third-party 3D-conv upscaler checkpoint separately; no separate upscaler node pack is required.

    Clean refiner branch

    The supplied workflow has two model paths:

    • Generation: shared attention patches → Turbo LoRA → Extra LoRA → optional Adaptive Native Cache.

    • Refinement: shared attention-patched H3 model before both LoRAs and cache.

    This keeps generation LoRAs and cache out of the refinement pass while retaining the shared attention patches.

    Recommended starting settings:

    • Refine steps: 5

    • Refine denoise: 0.25

    • Sampler: Euler

    • Scheduler: Simple

    • CFG: 1.0

    Five passes at 0.25 denoise sample the final part of a native 20-step trajectory. Other useful pairs are 4 / 0.20, 6 / 0.30 and 7 / 0.35.

    Native continuation and HQ export

    LongForge preserves the native video/audio latent for continuation and stores enhanced HQ video for previews and export. Audio is not upscaled and remains protected during refinement. The previous HQ overlap is carried into the next HQ scene.

    Keep the same HQ mode, checkpoint, precision and output geometry throughout the film. Regenerate from Scene 1 when changing these settings.

    For example, a film generated at 736 × 416 with a 2.0× HQ scale produces 1472 × 832 video while continuation still uses the original native latent. HQ increases processing time and storage requirements.

    Scene Studio

    Project & Scenes provides editable scene cards, search, generation statuses, saved previews, pending-scene reordering, regeneration, variation and project cleanup controls.

    Pending cards can be dragged into a new order while search is empty. Saved scenes depend on the latent context before them; changing an already-generated part requires rebuilding its continuation.

    Open project restores saved prompts and video, sampling and HQ settings. Model and LoRA selections remain in the workflow JSON. Save that JSON before closing the browser to retain pending editor changes and loader selections.

    Use a different project name for a separate film. New film resets the active history for the current name while retaining editor text and old takes. Clean unused removes takes outside the active chain; Delete film removes the entire project.

    Audio continuity

    Recommended starting settings:

    • Continuity mode: MANUAL

    • Overlap: 22 frames

    • Preserve audio overlap: enabled

    • Extra audio history: 0 frames

    You can select AUTO BALANCED or AUTO SAFE when you want LongForge to choose overlap from the preceding scene. A 5-frame overlap is supported but provides less context and a higher risk of motion or audio resets.

    Extra audio history sits immediately before the protected overlap. If active REF2VA audio references conflict with this additional history, LongForge disables only Extra audio history for that scene and reports it; protected audio overlap remains enabled.

    Long-audio references

    In REF2VA, choose Audio use → Follow scenes · REF2VA in the audio card. LongForge advances through the source according to the saved film timeline, avoiding a duplicate reference for audio already covered by the protected overlap.

    This guides H3’s own audio synthesis. It does not replace the final soundtrack or guarantee exact words, voices, rhythm or timing. Actual source intervals appear in the generation report. If the selected source ends early, the remaining reference window is padded with silence.

    Seam repair and diagnostic WAVs

    Repair audio clicks targets isolated clicks near joins. A broad dropout or sustained reset needs review and may require regeneration with more protected context.

    For an audible join, enable Diagnostic WAVs · export only, select EXPORT FILM and run again. No new diffusion generation is required. The export folder receives:

    • the MP4;

    • .before_repair.wav — continuous decoded H3 audio before seam repair;

    • .before_aac.wav — audio after seam repair and before AAC encoding.

    The WAVs retain float32 samples without lossy compression. Both reflect the Fade film edges setting. Compare files with the same export basename around the reported join.

    Repair video seams is optional and OFF by default. Fade film edges applies short ramps only at the beginning and end of the film.

    Recovery metadata

    Set Video metadata → RECOVERY to embed prompts and workflow recovery information in newly created previews and exported MP4s. This can include model/LoRA filenames and reference-media selections.

    Drag a recovery MP4 into ComfyUI to reopen its workflow. A scene preview restores that scene’s graph; a full-film export uses the last scene’s model setup with the active film prompts. It does not automatically switch a different LoRA setup for every card.

    Recovery does not contain model weights, LoRA files, original reference media or scene latents. Keep the project folder to continue generating after a restart:

    ComfyUI/output/longforge_native/<project>/

    The public workflow starts with Video metadata → NONE. Select RECOVERY when you want embedded restoration data. NONE affects newly written videos; it does not remove metadata from existing MP4s or delete project data. ComfyUI’s --disable-metadata also disables embedded recovery.

    If recovery export fails because of a JSON nan value, select NONE. This does not change generation itself.

    Project handling on Windows

    LongForge uses per-project locking, atomic JSON updates and protection against simultaneous generation and cleanup. Successfully generated scenes are saved immediately, including during ALL PENDING.

    Clean unused and Delete film require confirmation. Avoid editing or generating the same project simultaneously in several browser tabs.

    Existing compatible project/take folders can be reused. Install the supplied v3.7 workflow to access the current controls and clean refiner connection.

    Workflow structure and defaults

    The graph uses visible native model, CLIP, VAE and LoRA loaders; KJNodes attention patches; and LongForge project, continuity, reference, cache, HQ, prompt-assistant and export nodes.

    ComponentSupplied v3.7 settingPipelineFL2VAResolution1344 × 768Scene length158 frames at 24 FPSContinuityMANUAL, 22-frame overlapPreserve audio overlapONExtra audio history0Sampling20 steps, res_multistep, simple, CFG 1, denoise 1Turbo LoRA / Extra LoRABypassed; no files selectedFirst / Last frame loadersBypassed; no files selectedKJ attention patchesEnabledAdaptive Native CacheEnabled, BALANCEDHQ FinishOFFPrompt AssistantOptional; text model not connectedReference LibraryEmptyVideo metadataNONEDiagnostic WAVsOFF

    Sampling remains editable. Euler/simple is a straightforward starting point for native sampling and cache comparisons; the supplied JSON retains res_multistep as its selected sampler.

    Width and height must be multiples of 32. Scene length and overlap follow the H3 grid 5 + 17 × k frames, with overlap shorter than the next scene. There is no fixed 1 MP resolution cap; practical limits depend on the model and hardware.

    Scene duration includes overlap. Three 243-frame scenes with 22-frame overlap produce 243 + 221 + 221 = 685 unique frames, approximately 28.54 seconds at 24 FPS.

    Download

    Download both archives:

    • ComfyUI-H3-LongForge-v3.7-NodePack.zip — custom nodes, web interface, bundled prompt-writing guides and installation README.

    • H3_LongForge-v3.7-Workflow-Prompt-Guide.zip — H3_LongForge_FL2VA_REF2VA.json and the complete PROMPT_GUIDE_RU_EN.md user and prompt guide.

    Models, LoRAs, VAEs, Gemma weights and upscaler checkpoints are not included.

    LongForge v3.7 requires:

    • a compatible recent ComfyUI build with native MiniMax H3 support and the V3 node API;

    • a matching H3 FL2VA or REF2VA diffusion model;

    • a compatible H3 text encoder;

    • H3 video and audio VAEs;

    • a usable FFmpeg executable or the imageio-ffmpeg fallback.

    The supplied graph also uses ComfyUI-KJNodes and compatible SageAttention for its enabled attention patches. These patches are optional to LongForge’s core pipeline; install their dependencies to use the graph as supplied, or remove the patch nodes and reconnect the model path.

    Optional features require matching LoRAs, a compatible 3D latent-upscaler checkpoint, or a separate Gemma 4 text model for Prompt Assistant.

    MiniMax H3 models

    Comfy-Org / MiniMax-H3

    Alternative H3 text encoders

    INT8 ConvRot · NVFP4

    Choose a format compatible with your ComfyUI build and hardware. These encoders serve H3 conditioning; Prompt Assistant uses its own text-generation model.

    Optional Prompt Assistant model

    Comfy-Org / Gemma 4 text encoders

    The bundled guide lists gemma4_12b_int8_convrot.safetensors and gemma4_e4b_it_fp8_scaled.safetensors as compatible options for a ComfyUI build supporting them.

    Attention nodes

    ComfyUI-KJNodes

    Optional latent upscaler

    LBH-123-AI / MiniMax H3 Latent Upscaler

    Use a compatible minimax_h3_latent_upscaler_3d_conv_v1_*.safetensors checkpoint. This integration supports the 3D version, not the 2D checkpoint.

    Reference-media loading and upscaler integration are built into LongForge; neither requires a separate node pack.

    Installation — Windows Portable

    1. Close ComfyUI completely.

    2. Remove the previous ComfyUI-H3-LongForge-NodePack folder. Keep your projects under ComfyUI/output/longforge_native/.

    3. Extract the new node pack so this file exists: ComfyUI/custom_nodes/ComfyUI-H3-LongForge-NodePack/__init__.py.

    4. Do not merge versions or leave duplicate LongForge installations in custom_nodes.

    5. Install/update ComfyUI-KJNodes and compatible SageAttention for the supplied attention nodes.

    6. If your installation needs the FFmpeg fallback, run the following from ComfyUI_windows_portable:

    python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-H3-LongForge-NodePack\requirements.txt
    1. Restart ComfyUI and refresh the browser with Ctrl+F5.

    2. Open H3_LongForge_FL2VA_REF2VA.json from the workflow archive.

    3. Select the H3 model, H3 text encoder, video VAE and audio VAE. Match Pipeline to the model.

    4. Select optional LoRA or image files before enabling their loaders with Ctrl+B. Bypass Native Cache when using Turbo.

    Optional Prompt Assistant setup

    1. Use a ComfyUI build with native Gemma 4 and text-generation support.

    2. Download a compatible Gemma encoder into ComfyUI/models/text_encoders/ and restart ComfyUI.

    3. Add a native Load CLIP node in the Prompt Assistant area.

    4. Select the Gemma file and set type = stable_diffusion.

    5. Connect its CLIP output to Prompt Assistant → text_model.

    6. Write a short idea in a scene card, press Generate draft, review the result and choose Apply to scene.

    For several connected scenes, apply Scene 1’s draft before drafting Scene 2, then apply Scene 2 before drafting Scene 3. Each request receives the preceding card’s text. Review the actual generated scene separately when deciding how its continuation should develop.

    Optional HQ Finish setup

    1. Download a compatible 3D-conv latent-upscaler checkpoint.

    2. Place it in ComfyUI/models/latent_upscale_models/.

    3. Restart ComfyUI and select the checkpoint in HQ Finish.

    4. Choose UPSCALE or UPSCALE + REFINE.

    5. For refinement, retain the supplied clean refiner connection before both LoRAs and cache.

    Generate your film

    Give the project a unique name. Write one prompt per scene and use + Scene to add the next continuation.

    A simple manual prompt structure:

    VIDEO: Describe the subject, action, setting and camera.
    SOUND: Describe ambience, effects and dialogue.
    MUSIC: No music.

    These headings are optional. Full native H3 prompts, including the longer structure produced by Prompt Assistant, can also be used.

    Continuation prompts should describe the next stage of the same action while preserving the intended subject, setting, camera and sound. Do not put several repeated VIDEO sections in one card and expect separate scenes.

    REF2VA reference markers

    Use <Picture 1>, <Video 1> and <Audio 1> for available references. Optional aliases can make mixed inputs easier to address:

    hero=image1
    motion=video1
    score=audio1
    voice=video1_audio

    Then use {hero}, {motion}, {score} or {voice} in the prompt. Define only sources you actually loaded. With FIRST SCENE, later prompts should continue the generated scene without addressing static references that are no longer being supplied.

    Main actions

    Action Result

    GENERATE NEXT + ONE SCENE -> Generates the next pending scene.
    GENERATE NEXT + ALL PENDING -> Generates all remaining scenes in order, saving each one.
    REGENERATE SELECTED -> Rebuilds from the selected saved scene with its saved seed.
    REGENERATE LAST -> Rebuilds the last saved scene with its saved seed.
    VARY SELECTED -> Rebuilds the selected scene with its saved seed plus one.
    PREVIEW SELECTED -> Recreates a missing preview from saved latents without diffusion.
    EXPORT FILM -> Assembles the active saved scenes into one MP4 with generated audio.

    When regenerating, ONE SCENE replaces only the chosen scene and leaves its continuation pending. ALL PENDING also rebuilds the following scenes.

    If all cards are already saved, GENERATE NEXT regenerates the selected scene. Add a new card first when you want to extend the film. Export uses saved scenes and does not generate pending cards.

    Optional final post-processing

    For separate enhancement or frame interpolation after export:

    DLSS 5 Visual Enhancer · Downloads

    This is a separate application with its own requirements and is not required by LongForge.

    Description

    FAQ

    Workflows
    MiniMax H3

    Details

    Downloads
    113
    Platform
    CivitAI
    Platform Status
    Available
    Created
    10/1/2026
    Updated
    10/1/2026
    Deleted
    -

    Files

    minimaxH3LongforgeFL2VA_fl2vaREF2VAV40.zip