CivArchive
    [MiniMax H3] NSFW I2VA / T2VA / FL2VA / R2VA Workflows ๐Ÿ‘พ Qwen3.5 Auto Prompt | โšก MiniMax-H3 Turbo LoRA | ๐Ÿ”ˆ Native Audio | TensorRT Upscale | RIFE Interpolation - ๐Ÿ” OneClick - FL2VA Loop
    NSFW

    ๐ŸŸฃ Deploy on Runpod

    ๐ŸŸก Deploy on Vast.ai


    ComfyUI-QwenVL-Mod โ€” Enhanced Vision-Language with MiniMax H3 Version 2.8.0 (2026/09/02) โ€” ๐ŸŽฌ MiniMax H3 Native Video+Audio + NVFP4 Blackwell + SOL-ATTN + Spectrum (native only) + lightx2v Turbo LoRA + Wildcards + 10Eros-Max Support + Camera Tag Dropdown + Unified FL2VA Loop


    โฌ†๏ธ 2026/09/04 UPDATE โฌ†๏ธ

    ๐Ÿ”ง Turbo LoRA Switch โ€” larryvrh โ†’ lightx2v 8-step 768p

    All Turbo workflows now use lightx2v Turbo LoRA 8-step 768p (Apache-2.0) instead of larryvrh v4-600:

    • FL2VA: minimax_h3_fl2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors (~1.96 GB)

    • Ref2VA: minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors (~1.96 GB)

    • No custom node required โ€” standard ComfyUI LoRA loader works

    • Removed Larryvrh/ComfyUI-MiniMax-H3-Turbo custom node

    • Trained at 1344ร—768 โ€” matches our native resolution exactly

    ๐Ÿšซ Spectrum + Sol-Attn Bypass with Turbo โ€” Verified

    Community testing confirmed that Spectrum and Sol-Attn cause fallbacks and slowdown when used with Turbo LoRA at 8 steps:

    • Spectrum: forecasts steps that don't have enough context at 8-step โ€” fallbacks double the work. "Near-useless at 10 steps" (community)

    • Sol-Attn (tau 1.3โ†’0.8): aggressive sparse approximation on 8-step trajectory causes reconstruction fallbacks

    • Rule: bypass both Spectrum and Sol-Attn when Turbo LoRA is active. They are for 20-step native only

    Turbo pure: Turbo LoRA ON ยท Spectrum OFF ยท Sol-Attn OFF ยท euler + simple ยท 8 steps Native + Spectrum: Turbo LoRA OFF ยท Spectrum ON ยท Sol-Attn ON (tau 1.0) ยท res_multistep + simple ยท 20 steps

    ๐Ÿ”„ 10Eros-Max Switch โ€” DmitryDB โ†’ cicalooo TURBO Hybrid Beta3

    The 10Eros-Max model has been switched to cicalooo's ComfyUI-native INT8 ConvRot Turbo Hybrid Beta3:

    • Model: 10Eros_Max_h3_TURBO-hybrid_beta3_int8_convrot_skip_edges.safetensors (~22.5 GB)

    • TURBO fused in checkpoint โ€” no separate t8star LoRA needed

    • Boundary blocks 0, 1, 48, 49 left in BF16 for stability

    • Removed DmitryDB 10Eros model + t8star compatibility LoRA

    โšก Comfy Kitchen Attention โ€” Replaces SageAttention

    All MMH3 Docker/provisioning now uses --use-ck-attention (Comfy Kitchen Attention) instead of --use-sage-attention:

    • Faster on Blackwell (RTX 5090 / PRO 6000)

    • Better detail preservation

    • No separate SageAttention package needed

    โฌ†๏ธ 2026/09/02 UPDATE โฌ†๏ธ

    ๐ŸŽฅ Camera Tag Dropdown (19 movements)

    The QwenVL node now has a camera_tag dropdown โ€” no more typing [ORBIT] manually in the prompt. Select from 19 camera movements directly in the node UI:

    CategoryTagsStatic[STATIC_CAMERA], [LOCKED_OFF]Slow zoom[SLOW_ZOOM_IN], [SLOW_ZOOM_OUT]Fast zoom[FAST_ZOOM_IN], [FAST_ZOOM_OUT]Pan[PAN_LEFT], [PAN_RIGHT]Tilt[TILT_UP], [TILT_DOWN]Dolly[DOLLY_IN], [DOLLY_OUT]Tracking[TRACKING_LEFT], [TRACKING_RIGHT]Crane[CRANE_UP], [CRANE_DOWN]Other[ORBIT], [HANDHELD], [ROLL]

    How it works: when you select a tag, it's injected both at the start of the prompt and as a FINAL CAMERA DIRECTIVE at the end โ€” so Qwen 9B actually respects it despite recency bias on long prompts. The tag also gets a short description so Qwen knows exactly what to write.

    Subject stays alive: the directive explicitly tells Qwen that the camera tag controls ONLY the camera โ€” the subject must still have natural, lively action (breathing, gestures, expression, body motion) throughout the clip. No more "statue during orbit" problem.

    Manual tags still work: if you leave the dropdown on None and type [ORBIT] in your prompt, it's detected and injected automatically as a fallback.

    Available on all three QwenVL nodes: AILab_QwenVL, AILab_QwenVL_Advanced, and AILab_QwenVL_PromptEnhancer.

    ๐Ÿ”„ FL2VA Loop Merged into FL2VA โ€” One Workflow, Bypass Group

    The separate MiniMaxH3-Turbo-FL2VA-Loop-Qwen3.5 workflow is removed. The loop trim logic now lives inside the main FL2VA workflow, wrapped in a "Loop Trim" group that can be toggled via the rgthree Fast Groups Bypasser node:

    • Loop mode (trim active): the ImageFromBatch + ComfyMathExpression nodes trim the frozen tail (~5 frames) for seamless looping

    • Non-loop mode (trim bypassed): toggle the group off in the Bypasser โ†’ VAEDecode passes directly to RIFE/upscale, full frames preserved

    No more switching between two workflows โ€” just toggle the group.

    ๐Ÿงน PromptEnhancer Cleanup

    • Removed the redundant custom_system_prompt input โ€” enhancement_style (presets) + prompt_text (user input) cover all use cases

    • Removed CUSTOM_ONLY_STYLE ("โœ๏ธ Custom Only (no preset)") โ€” no longer needed

    • Added camera_tag dropdown (same as main QwenVL nodes)

    โฌ†๏ธ 2026/08/31 UPDATE โฌ†๏ธ

    ๐ŸŽฒ Wildcards (T2VA Workflow)

    The T2VA Turbo workflow includes a WildcardProcessor node that injects randomized prompt fragments from the MadBe's Prompt Engine (__mbe/prmpt/*) wildcard library. Each generation picks a random entry from each wildcard file, so the same seed produces different results across runs unless you pin the seed.

    Wildcards Used in the T2VA Workflow

    WildcardCategoryWhat it randomizes__mbe/prmpt/vidstyle__Video styleCinematic / anime / vintage film / 3D CG / claymation / watercolor / fantasy etc.__mbe/prmpt/imgcmp/shots__Camera shotClose-up / wide / medium / dolly / crane / orbit / handheld etc.__mbe/prmpt/lctns/rndmlctns__LocationRandom setting (bedroom / beach / studio / alley / forest / rooftop etc.)__mbe/prmpt/light/rndmlight__LightingKey light direction, color temperature, soft/hard, ambient mood__mbe/prmpt/char/favchar__CharacterRandom character archetype (age, body type, hair, ethnicity)__mbe/prmpt/clths/rndmclth__ClothingRandom outfit / garment description__mbe/prmpt/clths/nudty_brsts__NSFWBreast / nudity descriptors (NSFW preset)__mbe/actsolo/brstplng__NSFW actionSolo breast-play action verbs (NSFW preset)

    How It Works

    1. The WildcardProcessor node sits before the Qwen3-VL prompt enhancer

    2. At queue time, each __wildcard__ token is replaced with a random line from the corresponding .txt file inside ComfyUI/custom_nodes/ComfyUI-Wildcards/wildcards/mbe/prmpt/

    3. The expanded text is passed to Qwen3-VL, which converts it into the official MiniMax H3 prompt format

    4. Different seed = different wildcard picks โ€” use a fixed seed if you want reproducible results

    Customizing Wildcards

    • Edit existing: open the .txt files under ComfyUI/custom_nodes/ComfyUI-Wildcards/wildcards/mbe/prmpt/ and add/remove lines (one entry per line)

    • Add your own: create a new .txt file, e.g. mbe/prmpt/mytags.txt, then reference it as __mbe/prmpt/mytags__

    • Remove a wildcard: delete the __...__ token from the WildcardProcessor text field in the workflow

    • Disable randomization: replace the __wildcard__ token with a fixed string

    Required Custom Node

    The wildcard files ship with the custom node. If a wildcard resolves to empty, the custom node is missing or the wildcard folder is not installed.

    ๐Ÿ”„ Sampler Change โ€” MiniMaxH3TurboSampler โ†’ KSamplerSelect + MiniMaxH3SigmaShift

    All 4 Turbo workflows (T2VA, I2VA, FL2VA, R2VA) have been updated to use ComfyUI core nodes instead of the custom MiniMaxH3TurboSampler:

    • Removed: MiniMaxH3TurboSampler (custom node from Larryvrh/ComfyUI-MiniMax-H3-Turbo)

    • Added: KSamplerSelect (sampler: euler) + MiniMaxH3SigmaShift (shift_video=12, shift_audio=3) โ€” both ComfyUI core nodes, no custom node required

    • Scheduler: simple (unchanged)

    Why?

    • On ComfyUI v0.34.2+ with native ModelSamplingAV, the custom MiniMaxH3TurboSampler internally delegates to stock euler anyway โ€” the custom node is redundant

    • Using core nodes means the same workflow works with both:

      • Standard model (minimax_h3_fl2va_pruned_nvfp4_convrot_int8) + Turbo LoRA minimax_h3_turbo_v4_step600_ema

      • 10Eros-Max (10Eros_Max_H3_FL2VA-INT8-ConvRot-HQ) + T8 compatibility LoRA minimax_h3_fl2v_turbo_8step_v1.0_10ErosMax_beta1_pruned_compat_v001_T8

    • Just swap LoadDiffusionModel and LoraLoaderBypassModelOnly โ€” the sampler path stays the same

    10Eros-Max (Optional โ€” Experimental)

    • Model: 10Eros_Max_H3_FL2VA-INT8-ConvRot-HQ.safetensors (~23.5 GB) โ€” DmitryDB/MiniMax-H3-10Eros-Max-Quants

    • LoRA: minimax_h3_fl2v_turbo_8step_v1.0_10ErosMax_beta1_pruned_compat_v001_T8.safetensors (~1.96 GB) โ€” t8star/minimax_h3_turbo_4step_10ErosMax_test4_pruned_curveproj1025_T8

    • Sampler: euler + MiniMaxH3SigmaShift (shift 12/3) + simple scheduler โ€” same as standard model

    • โš ๏ธ NVFP4 degrades quality and audio on the 10Eros fine-tune โ€” use INT8 ConvRot HQ only

    • โš ๏ธ The T8 LoRA is checkpoint-specific โ€” only works with the exact 10Eros pruned model (SHA-256: f82cc3f723b080e7ae94a7c98f95aa989e387618d0bdc940133dfbd9f432c062)

    โฌ†๏ธ 2026/08/27 UPDATE โฌ†๏ธ

    NVFP4+INT8 ConvRot Hybrid โ€” New Default

    • Default diffusion models: NVFP4+INT8 ConvRot hybrid (minimax_h3_fl2va_pruned_nvfp4_convrot_int8 / minimax_h3_ref2va_pruned_nvfp4_convrot_int8) from lilcheaty/MiniMax-H3-NVFP4 โ€” NVFP4 on MLP, INT8 ConvRot on attention, BF16 on sensitive layers. Best speed/quality on Blackwell, ~2.5x faster than pure INT8 with identical visual quality

    • Uncensored text encoder: NVFP4 (qwen3vl_32b_heretic_minimax_h3_nvfp4) from Momoking/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4 โ€” fast on Blackwell, no quality impact on text encoding

    • โš ๏ธ Non-Blackwell GPUs: Use pure INT8 ConvRot models from Comfy-Org/MiniMax-H3 instead. NVFP4 requires sm_120+ (RTX 5090 / PRO 6000).

    SOL-ATTN + Spectrum Integration

    • All 4 Turbo workflows now include SOL-ATTN (Scheduled Sol Attention) for sharper output

    • All 4 Turbo workflows now include Spectrum adaptive smoothing (offline replay disabled for speed)

    • Turbo LoRA linked from preset to subgraph in all workflows

    Turbo Step Standardization

    • All Turbo workflows standardized to 8 steps (was 6 for T2VA/R2VA)

    • Consistent minimax_h3_turbo_v4_step600_ema LoRA across all workflows

    TensorRT Batch Size

    • RIFE and Upscaler TRT expose separate loader and runner batch_size parameters

    • Verified stable configuration:

      • RIFE: loader 1, runner 1

      • Upscaler: loader 2, runner 2

    • The loader compiles the TensorRT engine profile; the runner controls frames sent per infer() call (runner โ‰ค loader)

    • RIFE batch values above 1 can build but currently fail during interpolation; keep RIFE at 1/1

    • Upscaler 2/2 is verified on RTX PRO 6000 Blackwell; batch 4 fails to build with TensorRT 10.15 on sm_120

    • Changing the loader batch size requires a different engine; delete incompatible cached TRT engines before rebuilding

    Upscaler: Auto-detect Scale Factor

    • Removed the scale dropdown (2x/4x) from the Upscaler runner node โ€” it was redundant and error-prone

    • The loader now auto-detects the upscale factor from the model name (2x* โ†’ 2, 4x* โ†’ 4, x2plus โ†’ 2, x4plus โ†’ 4)

    • The factor is passed to the runner via the engine object โ€” no more mismatch between model and scale setting

    • Requires ComfyUI-Upscaler-TensorRT-Auto updated to latest version


    โš ๏ธ Requirements โ€” Read First!

    GPU & VRAM

    • ๐ŸŸข Recommended template configuration โ€” RTX 5090 (32 GB) / RTX PRO 6000 (48 GB) โ†’ NVFP4 diffusion + NVFP4 text encoder

    • ๐ŸŸก Non-Blackwell alternative โ€” RTX 4090 / 3090 (24 GB) โ†’ INT8 ConvRot + offload

    • ๐ŸŸ  Lower-VRAM alternative โ€” 12โ€“16 GB โ†’ INT4 + aggressive offload; slow and not recommended for production

    • โš ๏ธ NVFP4 requires Blackwell (sm_120+) and does not run on RTX 4090/3090/4080

    12 GB GPUs (e.g. RTX 3060 12GB): Technically possible with INT4 models + aggressive offloading, but very slow. You need 32 GB+ system RAM and a fast NVMe SSD. Not recommended for production use.

    Model Quantization Options

    Software

    • ComfyUI: v0.31.0+ (required for MiniMax H3 native support)

    • Python: 3.10+

    • CUDA: 12.8+ (13.0 recommended)

    • Storage: allow at least 90 GB for the complete provisioned package (~81 GB of models plus engines, workflows and outputs)

    Qwen3-VL Prompt Enhancer

    • GGUF: Q4_K_S (~4.8 GB) or Q5_K_S (~5.5 GB) for 8B model

    • HF: Qwen3-VL-8B-Heretic-Stable (~16 GB) or Qwen3-VL-4B (~8 GB)

    โšก MiniMax-H3 Turbo LoRA (Optional โ€” Faster & Sharper)

    A distilled 8-step LoRA for MiniMax-H3 that replaces the default ~20-step sampling, with a dedicated ComfyUI node. All Turbo workflows now include SOL-ATTN (Scheduled Sol Attention) and Spectrum adaptive smoothing for sharper, smoother output.

    Works with all tasks: T2VA, I2VA, FL2VA and R2VA.


    ๐ŸŒŸ What is ComfyUI-QwenVL-Mod?

    A powerful enhanced vision-language node for ComfyUI that combines Qwen3-VL models with MiniMax H3 video generation workflows. Features multilingual support, visual style detection, native stereo audio, and NSFW capabilities for professional AI content creation.

    Think: "Your all-in-one solution for intelligent prompt enhancement and video+audio generation with MiniMax H3!"


    ๐ŸŽฌ Key Features

    ๐Ÿš€ MiniMax H3 Video+Audio Generation

    • T2VA (Text-to-Video+Audio): Generate video with native stereo audio from text

    • I2VA (Image-to-Video+Audio): Animate a first-frame image with audio

    • FL2VA (First-Last-Frame): Generate the transition between two keyframes โ€” Qwen3-VL sees both frames

    • R2VA (Reference-to-Video): Lock character identity, style, motion, or voice using reference images

    ๐Ÿง  Qwen3-VL Auto-Prompting

    • Multilingual: Write your prompt in any language โ€” Qwen3-VL translates and converts it

    • Auto-format: Generates the official MiniMax H3 prompt format (3-field for base, 6-field for R2VA)

    • Multi-reference: Qwen3-VL sees all connected images via image + image2 inputs

    • Visual style detection: 12+ artistic styles (photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy, etc.)

    • Smart caching: Performance optimization with Fixed Seed Mode

    • GGUF backend: Efficient local model inference with quantization support

    • Qwen3.5 support: Thinking mode disabled via /no_think for fast prompt generation

    ๐Ÿ”Š Native Stereo Audio

    • No separate audio node needed โ€” MiniMax H3 generates video and audio jointly in a single forward pass

    • Voice, sound effects, and music modeled together, not layered on afterward

    • Describe sounds in your prompt and the model generates them natively

    ๐ŸŽจ NSFW Support

    • Comprehensive content generation without restrictions

    • 9 dedicated NSFW presets (3 base ๐ŸŽฌ + 3 R2VA ๐ŸŽž๏ธ + 3 FL2VA ๐Ÿ”„) with explicit diegetic soundscape

    • Natural progression, style adaptation, consistent characters


    ๐Ÿ“ฆ What's Included โ€” 4 Turbo Workflows

    All workflows are pre-wired with MiniMax-H3 Turbo LoRA + MiniMax-H3 Turbo Sampler at 8 steps + SOL-ATTN + Spectrum.

    1. โšก T2VA Turbo โ€” MiniMaxH3-Turbo-T2VA-Qwen3.5.json โ€” text only โ€” Text-to-video+audio. Simplest workflow.

    2. โšก I2VA Turbo โ€” MiniMaxH3-Turbo-I2VA-Qwen3.5.json โ€” text + first-frame image (image) โ€” Image-to-video. First-frame animation with audio.

    3. โšก FL2VA Turbo โ€” MiniMaxH3-Turbo-FL2VA-Qwen3.5.json โ€” text + first-frame (image) + last-frame (image2) โ€” First-Last-Frame to video. Includes TensorRT upscale + RIFE frame interpolation for 48 fps output.

    4. โšก R2VA Turbo โ€” MiniMaxH3-Turbo-R2VA-Qwen3.5.json โ€” text + reference images (image + image2) โ€” Reference-to-video. Lock identity, style, motion, camera, or voice using up to 9 ref images.

    Workflows 3 and 4 include TensorRT upscaling and RIFE frame interpolation for 48 fps high-resolution output.


    ๐Ÿ–ผ๏ธ Multi-Reference Input (image2)

    The QwenVL-Mod node has two image inputs:

    • T2VA: no images needed

    • I2VA: image = first frame

    • FL2VA: image = first frame, image2 = last frame, frame_count = 1

    • R2VA: image = primary reference, image2 = additional references (batch, up to 9), frame_count = 1โ€“9

    Qwen3-VL sees all connected images as individual images (not as a video sequence), enabling proper multi-reference analysis for FL2VA and R2VA.


    ๐ŸŽฏ QwenVL-Mod NSFW Presets (9 total)

    The workflows include built-in NSFW presets for the Qwen3-VL prompt enhancer:

    ๐ŸŽฌ Base Presets (T2VA / I2VA)

    • ๐ŸŽฌ MiniMax H3 NSFW (5s) โ€” 5 seconds โ€” 3 fields: integrated_multimodal_description + overall_soundscape + non_diegetic_music

    • ๐ŸŽฌ MiniMax H3 NSFW (10s) โ€” 10 seconds โ€” Same format

    • ๐ŸŽฌ MiniMax H3 NSFW (15s) โ€” 15 seconds โ€” Same format

    ๐Ÿ”„ FL2VA Presets (First-Last-Frame)

    • ๐Ÿ”„ MiniMax H3 NSFW FL2VA (5s) โ€” 5 seconds โ€” 3 fields, transition-focused (describes the path between frames)

    • ๐Ÿ”„ MiniMax H3 NSFW FL2VA (10s) โ€” 10 seconds โ€” Same format

    • ๐Ÿ”„ MiniMax H3 NSFW FL2VA (15s) โ€” 15 seconds โ€” Same format

    ๐ŸŽž๏ธ R2VA Presets (Reference)

    • ๐ŸŽž๏ธ MiniMax H3 NSFW R2VA (5s) โ€” 5 seconds โ€” 6 fields: subject_definitions + summary + retention_analysis + detailed_description + overall_soundscape + non_diegetic_music

    • ๐ŸŽž๏ธ MiniMax H3 NSFW R2VA (10s) โ€” 10 seconds โ€” Same format

    • ๐ŸŽž๏ธ MiniMax H3 NSFW R2VA (15s) โ€” 15 seconds โ€” Same format

    What the presets produce

    • ๐ŸŽฌ Base: [Shot 1] with style + initial composition, camera vocabulary, speaker IDs, diegetic soundscape

    • ๐Ÿ”„ FL2VA: Describes the transition path between first and last frames (not the scene โ€” images fix the scene). Favors single continuous shot.

    • ๐ŸŽž๏ธ R2VA: 6-section format with <Subject N>, <Picture N>, <Video N>, <Audio N> labels, retention markers (fully_preserved, partially_preserved, etc.), task-type summary

    • All presets: smooth, continuous camera motion (no abrupt or stepped changes), explicit diegetic soundscape, optional non-diegetic music (defaults to N/A)

    SFW presets are also available. Edit the preset dropdown in the QwenVL node to switch.


    ๐ŸŽฎ Usage Examples

    Basic Text-to-Video (T2VA)

    1. Load MiniMaxH3-Turbo-T2VA-Qwen3.5.json

    2. Write your prompt in any language

    3. Select preset ๐ŸŽฌ MiniMax H3 NSFW (5s/10s/15s)

    4. Generate video with native audio

    Image-to-Video (I2VA)

    1. Load MiniMaxH3-Turbo-I2VA-Qwen3.5.json

    2. Upload your first-frame image to image

    3. Select preset ๐ŸŽฌ MiniMax H3 NSFW (5s/10s/15s)

    4. Write what happens next (in any language)

    5. Generate animated video with audio

    First-Last-Frame (FL2VA)

    1. Load MiniMaxH3-Turbo-FL2VA-Qwen3.5.json

    2. Upload first-frame to image, last-frame to image2, set frame_count=1

    3. Select preset ๐Ÿ”„ MiniMax H3 NSFW FL2VA (5s/10s/15s)

    4. Describe the transition between the two frames

    5. Generate the interpolated video at 48 fps with TensorRT upscale + RIFE

    Reference-to-Video (R2VA)

    1. Load MiniMaxH3-Turbo-R2VA-Qwen3.5.json

    2. Upload primary reference to image, additional references to image2 (batch), set frame_count to match

    3. Select preset ๐ŸŽž๏ธ MiniMax H3 NSFW R2VA (5s/10s/15s)

    4. Reference them by tag in your prompt: <Picture 1>, <Picture 2>, etc.

    5. Generate video with locked identity/style


    ๐Ÿ”ง Technical Specifications

    โšก Performance

    • Output: 768p, 24 fps (native), up to ~15 seconds

    • Audio: Native stereo, generated jointly with video

    • Upscale: TensorRT RealESRGAN x4 (FL2VA + R2VA workflows)

    • Frame interpolation: RIFE v4.25 โ†’ 48 fps (FL2VA + R2VA workflows)

    • Sage Attention: FP16 accumulation, async offload

    • Smart caching: Reuse prompts with same inputs, Fixed Seed Mode for text-only caching

    ๐ŸŽจ Model Support

    • Qwen3-VL 4B: 7 GGUF variants (2.38 GB โ€“ 4.28 GB)

    • Qwen3-VL 8B: 7 GGUF variants (4.8 GB โ€“ 8.71 GB)

    • Qwen3.5: 4B / 9B / 27B (uncensored, heretic, unsloth) โ€” thinking mode disabled

    • HF Models: Josiefed, official, Heretic-Stable variants

    • Quantization: Q4_K_S, Q5_K_S, FP16, INT8

    ๐ŸŒ Multilingual Capabilities

    • Input languages: Any language supported

    • Auto-translation: Automatic translation to optimized English

    • Style detection: Works with multilingual prompts

    • Cultural adaptation: Context-aware prompt enhancement


    ๐Ÿ“ฆ Installation

    Quick Install

    1. Download: ComfyUI-QwenVL-Mod (latest version)

    2. Extract to ComfyUI/custom_nodes/ComfyUI-QwenVL-Mod

    3. Install requirements: pip install -r requirements.txt

    4. Restart ComfyUI

    5. Load included workflows from minimax/ folder

    Custom Nodes Required

    Note: ComfyMathExpression is built into ComfyUI core (v0.24.1+) โ€” no custom node needed.

    Models Required

    T2VA / I2VA / FL2VA use the NVFP4 FL2VA model (Blackwell GPUs):

    • models/vae/ โ†’ minimax_h3_video_vae_fp16.safetensors (~5 GB)

    • models/vae/ โ†’ minimax_h3_audio_vae_fp32.safetensors (~0.6 GB)

    • models/diffusion_models/ โ†’ minimax_h3_fl2va_pruned_nvfp4_convrot_int8.safetensors (~20 GB) โ€” lilcheaty/MiniMax-H3-NVFP4

    • models/text_encoders/ โ†’ qwen3vl_32b_heretic_minimax_h3_nvfp4.safetensors (~15.7 GB) โ€” Momoking/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4

    R2VA (ref2va) uses the NVFP4 Ref2VA model:

    • models/diffusion_models/ โ†’ minimax_h3_ref2va_pruned_nvfp4_convrot_int8.safetensors (~20 GB) โ€” lilcheaty/MiniMax-H3-NVFP4

    Non-Blackwell GPUs: Use INT8 ConvRot models from Comfy-Org/MiniMax-H3 instead. NVFP4 requires sm_120+ (RTX 5090 / PRO 6000).

    10Eros-Max NVFP4 (optional/experimental)

    • models/diffusion_models/ โ†’ 10Eros_Max_H3_FL2VA-INT8-ConvRot-HQ.safetensors (~23.5 GB)

    • Pair it only with minimax_h3_fl2v_turbo_8step_v1.0_10ErosMax_beta1_pruned_compat_v001_T8.safetensors (~1.96 GB)

    • Switch both diffusion model and matching LoRA together; do not mix the standard and 10Eros LoRAs

    • The standard MiniMax H3 NVFP4+INT8 hybrid is the verified default. 10Eros-Max remains experimental; NVFP4 HQ degrades quality and audio on the 10Eros fine-tune, so INT8 ConvRot HQ is used instead

    Turbo LoRA (standard โ€” all tasks)

    INT4 alternative (for 12-16 GB GPUs): Merserk/MiniMax-H3-INT4-ConvRot

    Qwen3-VL Prompt Enhancer

    • models/LLM/ โ†’ Qwen3-VL-8B-Heretic-Stable (GGUF or HF)

    TensorRT Engines (FL2VA + R2VA only)

    • models/upscale_models/ โ†’ RealESRGAN_x4 (TensorRT engine)

    • models/rife/ โ†’ rife425_ensemble_False_scale_1_sim (TensorRT engine, ONNX auto-downloaded from HF)

    TensorRT engines must be built for your specific GPU. See ComfyUI-RIFE-TensorRT-Auto and ComfyUI-Upscaler-TensorRT-Auto for build instructions.


    ๐ŸŽฌ MiniMax H3 Prompting Notes

    How to Write Your Prompt

    Describe the scene naturally. Be clear about the concepts below โ€” Qwen3-VL handles the rest:

    • ๐ŸŽจ Visual style (put it first): photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy

    • ๐Ÿ‘ฅ Subjects: number, gender, appearance, clothing, position, expression

    • ๐Ÿƒ Action / motion: what happens, speed, interaction

    • ๐ŸŽฅ Camera: dolly, pan, zoom, static, handheld, crane, orbit โ€” smooth and continuous (no abrupt changes)

    • ๐ŸŒ Environment: setting, lighting, atmosphere, time of day

    • ๐Ÿ”Š Audio (important!): dialogue, breaths, moans, skin contact, ambient sounds, music

    ๐Ÿ”„ FL2VA: Describe the transition between frames, not the scene (images fix the scene) ๐ŸŽž๏ธ R2VA: Reference inputs by tag: <Picture 1>, <Picture 2>, <Video 1>, <Audio 1>

    Resolution Guidance

    MiniMax H3 native canvas: 768 px short edge, long edge capped at 1344 px, multiples of 32.

    • ๐Ÿ“ฑ Portrait: 768ร—1344 ยท 896ร—1152 ยท 960ร—1280

    • โฌ› Square: 1024ร—1024

    • ๐Ÿ–ฅ๏ธ Landscape: 1344ร—768 ยท 1152ร—896 ยท 1280ร—960

    โš ๏ธ Match the aspect ratio to your input image! Forcing 16:9 on a portrait image will squash it.

    โš ๏ธ Avoid direct 1080p. Generate at native resolution, then upscale with TensorRT nodes (FL2VA + R2VA workflows).

    Duration

    Choose a preset: 5s / 10s / 15s. The Math Expression node snaps the frame count to the model's 17-frame-per-block grid (17k+5 at 24 fps).


    ๐Ÿณ Docker / Cloud Ready

    OneClick RunPod Template

    Prefer a ready-to-go environment? Use the OneClick - ComfyUI - MiniMax H3 Turbo - Qwen3VL RunPod template:

    • Docker image: huchukato/comfyui-qwenvl-runpod:cu13-mmh3

    • Base: huchukato/comfyui-base:cu130

    • All custom nodes pre-installed

    • All 4 Turbo workflows auto-downloaded at boot

    • Models auto-downloaded at first boot (~81 GB including INT8 diffusion, NVFP4 text encoder, 10Eros INT8 HQ and Turbo LoRAs; persistent)

    • ComfyUI v0.34.2 baked into base image

    • Sage Attention, FP16 accumulation, async offload

    • TensorRT upscaling + RIFE interpolation (stable defaults: Upscaler 2/2, RIFE 1/1)

    • SOL-ATTN + Spectrum for Turbo workflows

    Access: ComfyUI :8188 ยท JupyterLab :8888 ยท FileBrowser :8080 (user admin / password adminadmin12) ยท SSH ssh root@pod-ip

    ๐Ÿ“– README & instructions

    ComfyUI Args (pre-configured)

    --disable-auto-launch
    --fast fp16_accumulation
    --use-sage-attention
    --reserve-vram 2
    --cuda-malloc
    --async-offload
    

    ๐Ÿš€ Why Choose ComfyUI-QwenVL-Mod + MiniMax H3?

    ๐ŸŽฌ For Content Creators

    • Native audio: Video and audio in one pass โ€” no separate MMAudio needed

    • Multilingual: Write in any language, Qwen3-VL handles translation

    • Professional: Official MiniMax H3 prompt format with camera vocabulary and speaker tags

    • Quality: 768p native, TensorRT upscale to higher resolution

    ๐Ÿ”ฅ For NSFW Content

    • Explicit: Uncensored generation with dedicated NSFW presets

    • 9 presets: 3 base ๐ŸŽฌ + 3 FL2VA ๐Ÿ”„ + 3 R2VA ๐ŸŽž๏ธ โ€” each tuned for its mode

    • Detailed: Rich scene descriptions with explicit diegetic soundscape

    • Natural: Realistic progression, consistent characters

    • Audio: Native moans, breaths, skin contact, ambient sounds

    โšก For Power Users

    • Customizable: Easy to modify presets and system prompts

    • Extendable: Add your own Qwen3-VL models (GGUF or HF)

    • Integrable: Works with existing ComfyUI setups

    • Optimized: Sage Attention, FP16, async offload, smart caching

    • Multi-reference: image2 input for FL2VA and R2VA workflows


    ๐ŸŒŸ What Makes This Special?

    • First: Complete MiniMax H3 workflow pack with Qwen3-VL auto-prompting

    • Native audio: No separate audio node โ€” MiniMax H3 does it all

    • 4 Turbo workflows: T2VA, I2VA, FL2VA, R2VA โ€” covers all MiniMax H3 modes

    • Multi-reference: Qwen3-VL sees all connected images (not just the first)

    • TensorRT: Built-in upscaling and frame interpolation

    • 9 NSFW presets: Dedicated presets for each mode with correct prompt structure

    • Multilingual: Any input language, auto-translated and formatted

    • Ready: Works out-of-the-box with included workflows


    ๐ŸŽฏ What's New in v2.6.0

    โšก NVFP4+INT8 ConvRot Hybrid โ€” New Default

    • โœ… Default diffusion models: NVFP4+INT8 ConvRot hybrid (minimax_h3_fl2va_pruned_nvfp4_convrot_int8 / minimax_h3_ref2va_pruned_nvfp4_convrot_int8)

    • โœ… ~2.5x faster than pure INT8 ConvRot with identical visual quality on Blackwell

    • โœ… NVFP4 on MLP, INT8 ConvRot on attention, BF16 on sensitive layers (by rockerBOO/lilcheaty)

    • โœ… NVFP4 requires Blackwell GPUs (RTX 5090 / PRO 6000, sm_120+)

    • โœ… Pure INT8 ConvRot models still available for non-Blackwell GPUs

    ๐Ÿ”ง SOL-ATTN + Spectrum Integration

    • โœ… All 4 Turbo workflows now include SOL-ATTN (Scheduled Sol Attention) for sharper output

    • โœ… All 4 Turbo workflows now include Spectrum adaptive smoothing (offline replay disabled for speed)

    • โœ… Turbo LoRA linked from preset to subgraph in all workflows

    โšก Turbo Step Standardization

    • โœ… All Turbo workflows standardized to 8 steps (was 6 for T2VA/R2VA)

    • โœ… Consistent minimax_h3_turbo_v4_step600_ema LoRA across all workflows

    ๐Ÿ“ฆ TensorRT Batch Size

    • โœ… Stable defaults: RIFE loader/runner 1/1, Upscaler loader/runner 2/2

    • โœ… Upscaler 2/2 verified on RTX PRO 6000 Blackwell

    • โš ๏ธ RIFE batch >1 currently fails during interpolation even when the engine builds; keep it at 1/1

    • โš ๏ธ Upscaler batch 4 fails to build with TensorRT 10.15 on Blackwell sm_120

    ๐Ÿงน Cleanup

    • โœ… Removed was-node-suite from MiniMax Dockerfile (not used by MiniMax workflows)

    • โœ… Removed ComfyUI-Frame-Interpolation from MiniMax (uses RIFE TensorRT instead)

    • โœ… Removed KJNodes from provisioning (baked into RunPod base image)

    • โœ… ComfyMathExpression is built into ComfyUI core โ€” no custom node needed

    ๐Ÿง  Qwen3.5 Thinking Fix

    • โœ… /no_think prefix for Qwen3.5 models (enable_thinking deprecated in recent llama.cpp)

    • โœ… Broadened architecture detection (qwen35, qwen35moe, qwen35_vl)

    • โœ… Works across both HF and GGUF nodes

    ๐Ÿ“ฆ Workflow Organization

    • โœ… Moved workflows to minimax/ folder

    • โœ… Renamed FLF to FL2VA (clearer naming)

    • โœ… Added Civitai documentation


    ๐Ÿ“‹ Credits


    ๐Ÿ“„ License

    Workflows are released under the same license as the underlying models and custom nodes. See each repository for details.

    MiniMax H3 model weights: Comfy-Org/MiniMax-H3 โ€” MiniMax H3 Community License.


    Built with โค๏ธ for the ComfyUI community

    Description

    FAQ

    ComfyWorkflows
    MiniMax H3

    Details

    Downloads
    36
    Platform
    CivitAI
    Platform Status
    Available
    Created
    9/8/2026
    Updated
    9/11/2026
    Deleted
    -

    Files

    MinimaxH3NSFWI2VAT2VAFL2VAR2VAWorkflowsQwen35_OneclickFL2VALoop.zip