ComfyUI-QwenVL-Mod
Enhanced Vision-Language with MiniMax H3 Version 2.5 (2026/08/05)
๐ฌ MiniMax H3 Native Video+Audio + Qwen3-VL Auto-Prompting
๐ฃ Deploy on Runpod

๐ก Deploy on Vast.ai

โ ๏ธ Requirements โ Read First!
GPU & VRAM
๐ข Recommended โ RTX 5090 / 4090 (24 GB+) โ INT8 pruned โ Fast, best quality
๐ก Enthusiast โ RTX 3090 / 4080 (16-24 GB) โ INT8 pruned + offload โ Slower, good quality
๐ Budget โ RTX 3060 / 4060 Ti (12-16 GB) โ INT4 pruned + offload โ Slow, usable quality
๐ด Experimental โ Blackwell GPUs (12 GB+) โ NVFP4 โ Requires Blackwell Tensor Cores
12 GB GPUs (e.g. RTX 3060 12GB): Technically possible with INT4 models + aggressive offloading, but very slow. You need 32 GB+ system RAM and a fast NVMe SSD. Not recommended for production use.
Model Quantization Options
BF16 (full) โ Diffusion ~42 GB + Text encoder ~65 GB = ~110 GB total โ Comfy-Org/MiniMax-H3
INT8 (pruned) โ Diffusion ~21 GB + Text encoder ~27 GB = ~50 GB total โ Comfy-Org/MiniMax-H3
INT4 (pruned) โ Diffusion ~11 GB + Text encoder ~15 GB = ~27 GB total โ Merserk/MiniMax-H3-INT4-ConvRot
NVFP4 (pruned) โ Diffusion ~12.5 GB + Text encoder ~15 GB = ~28 GB total โ Comfy-Org/MiniMax-H3 (Blackwell)
Software
ComfyUI: v0.30.0+ (required for MiniMax H3 native support)
Python: 3.10+
CUDA: 12.8+ (13.0 recommended)
Storage: 30-110 GB SSD depending on quantization
Qwen3-VL Prompt Enhancer
GGUF: Q4_K_S (~4.8 GB) or Q5_K_S (~5.5 GB) for 8B model
HF: Qwen3-VL-8B-Heretic-Stable (~16 GB) or Qwen3-VL-4B (~8 GB)
๐ What is ComfyUI-QwenVL-Mod?
A powerful enhanced vision-language node for ComfyUI that combines Qwen3-VL models with MiniMax H3 video generation workflows. Features multilingual support, visual style detection, native stereo audio, and NSFW capabilities for professional AI content creation.
Think: "Your all-in-one solution for intelligent prompt enhancement and video+audio generation with MiniMax H3!"
๐ฌ Key Features
๐ MiniMax H3 Video+Audio Generation
T2VA (Text-to-Video+Audio): Generate video with native stereo audio from text
I2VA (Image-to-Video+Audio): Animate a first-frame image with audio
FL2VA (First-Last-Frame): Generate the transition between two keyframes โ Qwen3-VL sees both frames
R2VA (Reference-to-Video): Lock character identity, style, motion, or voice using reference images
๐ง Qwen3-VL Auto-Prompting
Multilingual: Write your prompt in any language โ Qwen3-VL translates and converts it
Auto-format: Generates the official MiniMax H3 prompt format (3-field for base, 6-field for R2VA)
Multi-reference: Qwen3-VL sees all connected images via
image+image2inputsVisual style detection: 12+ artistic styles (photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy, etc.)
Smart caching: Performance optimization with Fixed Seed Mode
GGUF backend: Efficient local model inference with quantization support
Qwen3.5 support: Thinking mode disabled via
/no_thinkfor fast prompt generation
๐ Native Stereo Audio
No separate audio node needed โ MiniMax H3 generates video and audio jointly in a single forward pass
Voice, sound effects, and music modeled together, not layered on afterward
Describe sounds in your prompt and the model generates them natively
๐จ NSFW Support
Comprehensive content generation without restrictions
9 dedicated NSFW presets (3 base ๐ฌ + 3 R2VA ๐๏ธ + 3 FL2VA ๐) with explicit diegetic soundscape
Natural progression, style adaptation, consistent characters
๐ฆ What's Included โ 4 Workflows
๐ T2VA โ
MiniMaxH3-T2VA-Qwen3VL.jsonโ text only โ Text-to-video+audio. Simplest workflow. Uses PromptEnhancer (text-only).๐ผ๏ธ I2VA โ
MiniMaxH3-I2VA-Qwen3VL.jsonโ text + first-frame image (image) โ Image-to-video. First-frame animation with audio.๐ FL2VA โ
MiniMaxH3-FL2VA-Qwen3VL.jsonโ text + first-frame (image) + last-frame (image2) โ First-Last-Frame to video. Qwen3-VL sees both frames and describes the transition. Includes TensorRT upscale + RIFE frame interpolation for 48 fps output.๐๏ธ R2VA โ
MiniMaxH3-R2VA-Qwen3VL.jsonโ text + reference images (image+image2) โ Reference-to-video. Qwen3-VL sees all references. Lock identity, style, motion, camera, or voice using up to 9 ref images.
Workflows 3 and 4 include TensorRT upscaling (RealESRGAN x4) and RIFE frame interpolation (rife49) for 48 fps high-resolution output.
๐ผ๏ธ Multi-Reference Input (image2)
The QwenVL-Mod node has two image inputs:
T2VA: no images needed
I2VA:
image= first frameFL2VA:
image= first frame,image2= last frame,frame_count= 1R2VA:
image= primary reference,image2= additional references (batch, up to 9),frame_count= 1โ9
Qwen3-VL sees all connected images as individual images (not as a video sequence), enabling proper multi-reference analysis for FL2VA and R2VA.
๐ฏ QwenVL-Mod NSFW Presets (9 total)
The workflows include built-in NSFW presets for the Qwen3-VL prompt enhancer:
๐ฌ Base Presets (T2VA / I2VA)
๐ฌ MiniMax H3 NSFW (5s)โ 5 seconds โ 3 fields:integrated_multimodal_description+overall_soundscape+non_diegetic_music๐ฌ MiniMax H3 NSFW (10s)โ 10 seconds โ Same format๐ฌ MiniMax H3 NSFW (15s)โ 15 seconds โ Same format
๐ FL2VA Presets (First-Last-Frame)
๐ MiniMax H3 NSFW FL2VA (5s)โ 5 seconds โ 3 fields, transition-focused (describes the path between frames)๐ MiniMax H3 NSFW FL2VA (10s)โ 10 seconds โ Same format๐ MiniMax H3 NSFW FL2VA (15s)โ 15 seconds โ Same format
๐๏ธ R2VA Presets (Reference)
๐๏ธ MiniMax H3 NSFW R2VA (5s)โ 5 seconds โ 6 fields:subject_definitions+summary+retention_analysis+detailed_description+overall_soundscape+non_diegetic_music๐๏ธ MiniMax H3 NSFW R2VA (10s)โ 10 seconds โ Same format๐๏ธ MiniMax H3 NSFW R2VA (15s)โ 15 seconds โ Same format
What the presets produce
๐ฌ Base:
[Shot 1]with style + initial composition, camera vocabulary, speaker IDs, diegetic soundscape๐ FL2VA: Describes the transition path between first and last frames (not the scene โ images fix the scene). Favors single continuous shot.
๐๏ธ R2VA: 6-section format with
<Subject N>,<Picture N>,<Video N>,<Audio N>labels, retention markers (fully_preserved,partially_preserved, etc.), task-type summaryAll presets: smooth, continuous camera motion (no abrupt or stepped changes), explicit diegetic soundscape, optional non-diegetic music (defaults to N/A)
SFW presets are also available. Edit the preset dropdown in the QwenVL node to switch.
๐ฎ Usage Examples
Basic Text-to-Video (T2VA)
Load
MiniMaxH3-T2VA-Qwen3VL.jsonWrite your prompt in any language
Select preset
๐ฌ MiniMax H3 NSFW (5s/10s/15s)Generate video with native audio
Image-to-Video (I2VA)
Load
MiniMaxH3-I2VA-Qwen3VL.jsonUpload your first-frame image to
imageSelect preset
๐ฌ MiniMax H3 NSFW (5s/10s/15s)Write what happens next (in any language)
Generate animated video with audio
First-Last-Frame (FL2VA)
Load
MiniMaxH3-FL2VA-Qwen3VL.jsonUpload first-frame to
image, last-frame toimage2, setframe_count=1Select preset
๐ MiniMax H3 NSFW FL2VA (5s/10s/15s)Describe the transition between the two frames
Generate the interpolated video at 48 fps with TensorRT upscale + RIFE
Reference-to-Video (R2VA)
Load
MiniMaxH3-R2VA-Qwen3VL.jsonUpload primary reference to
image, additional references toimage2(batch), setframe_countto matchSelect preset
๐๏ธ MiniMax H3 NSFW R2VA (5s/10s/15s)Reference them by tag in your prompt:
<Picture 1>,<Picture 2>, etc.Generate video with locked identity/style
๐ง Technical Specifications
โก Performance
Output: 768p, 24 fps (native), up to ~15 seconds
Audio: Native stereo, generated jointly with video
Upscale: TensorRT RealESRGAN x4 (FL2VA + R2VA workflows)
Frame interpolation: RIFE rife49 โ 48 fps (FL2VA + R2VA workflows)
Sage Attention: FP16 accumulation, async offload
Smart caching: Reuse prompts with same inputs, Fixed Seed Mode for text-only caching
๐จ Model Support
Qwen3-VL 4B: 7 GGUF variants (2.38 GB โ 4.28 GB)
Qwen3-VL 8B: 7 GGUF variants (4.8 GB โ 8.71 GB)
Qwen3.5: 4B / 9B / 27B (uncensored, heretic, unsloth) โ thinking mode disabled
HF Models: Josiefed, official, Heretic-Stable variants
Quantization: Q4_K_S, Q5_K_S, FP16, INT8
๐ Multilingual Capabilities
Input languages: Any language supported
Auto-translation: Automatic translation to optimized English
Style detection: Works with multilingual prompts
Cultural adaptation: Context-aware prompt enhancement
๐ฆ Installation
Quick Install
Download: ComfyUI-QwenVL-Mod (latest version)
Extract to
ComfyUI/custom_nodes/ComfyUI-QwenVL-ModInstall requirements:
pip install -r requirements.txtRestart ComfyUI
Load included workflows from
minimax/folder
Custom Nodes Required
ComfyUI-QwenVL-Mod โ All workflows (Qwen3-VL prompt enhancer) โ huchukato/ComfyUI-QwenVL-Mod
ComfyUI-RIFE-TensorRT-Auto โ FL2VA, R2VA (frame interpolation) โ huchukato/ComfyUI-RIFE-TensorRT-Auto
ComfyUI-Upscaler-TensorRT-Auto โ FL2VA, R2VA (upscaling) โ huchukato/ComfyUI-Upscaler-TensorRT-Auto
ComfyUI-VideoHelperSuite โ FL2VA, R2VA (VHS_VideoCombine) โ Kosinkadink/ComfyUI-VideoHelperSuite
ComfyUI-Easy-Use โ FL2VA, R2VA (easy showAnything) โ yolain/ComfyUI-Easy-Use
comfyui-find-perfect-resolution โ All workflows (ResolutionSelector) โ ashtar1984/comfyui-find-perfect-resolution
was-node-suite-comfyui โ R2VA (ComfyMathExpression) โ ltdrdata/was-node-suite-comfyui
Models Required
All MiniMax H3 models from Comfy-Org/MiniMax-H3 on Hugging Face.
T2VA / I2VA / FL2VA (fl2va)
models/vae/โminimax_h3_video_vae_fp16.safetensors(~5 GB)models/vae/โminimax_h3_audio_vae_fp32.safetensors(~0.6 GB)models/diffusion_models/โminimax_h3_fl2va_pruned_int8_convrot.safetensors(~21 GB)models/text_encoders/โqwen3vl_32b_minimax_h3_int8_convrot.safetensors(~27 GB)
R2VA (ref2va) โ same as above, except:
models/diffusion_models/โminimax_h3_ref2va_pruned_int8_convrot.safetensors(~21 GB)
INT4 alternative (for 12-16 GB GPUs): Merserk/MiniMax-H3-INT4-ConvRot
Qwen3-VL Prompt Enhancer
models/LLM/โQwen3-VL-8B-Heretic-Stable(GGUF or HF)
TensorRT Engines (FL2VA + R2VA only)
models/upscale_models/โRealESRGAN_x4(TensorRT engine)models/rife/โrife49_ensemble_True_scale_1_sim(TensorRT engine)
TensorRT engines must be built for your specific GPU. See ComfyUI-RIFE-TensorRT-Auto and ComfyUI-Upscaler-TensorRT-Auto for build instructions.
Download Links
VAE: video_vae_fp16 ยท audio_vae_fp32
Diffusion (fl2va): minimax_h3_fl2va_pruned_int8_convrot.safetensors
Diffusion (ref2va): minimax_h3_ref2va_pruned_int8_convrot.safetensors
Text encoder: qwen3vl_32b_minimax_h3_int8_convrot.safetensors
INT4 models: Merserk/MiniMax-H3-INT4-ConvRot
๐ฌ MiniMax H3 Prompting Notes
How to Write Your Prompt
Describe the scene naturally. Be clear about the concepts below โ Qwen3-VL handles the rest:
๐จ Visual style (put it first):
photorealistic,cinematic,anime,3D CG,claymation,vintage film,watercolor,fantasy๐ฅ Subjects: number, gender, appearance, clothing, position, expression
๐ Action / motion: what happens, speed, interaction
๐ฅ Camera: dolly, pan, zoom, static, handheld, crane, orbit โ smooth and continuous (no abrupt changes)
๐ Environment: setting, lighting, atmosphere, time of day
๐ Audio (important!): dialogue, breaths, moans, skin contact, ambient sounds, music
๐ FL2VA: Describe the transition between frames, not the scene (images fix the scene) ๐๏ธ R2VA: Reference inputs by tag:
<Picture 1>,<Picture 2>,<Video 1>,<Audio 1>
Resolution Guidance
MiniMax H3 native canvas: 768 px short edge, long edge capped at 1344 px, multiples of 32.
๐ฑ Portrait: 768ร1344 ยท 896ร1152 ยท 960ร1280
โฌ Square: 1024ร1024
๐ฅ๏ธ Landscape: 1344ร768 ยท 1152ร896 ยท 1280ร960
โ ๏ธ Match the aspect ratio to your input image! Forcing 16:9 on a portrait image will squash it.
โ ๏ธ Avoid direct 1080p. Generate at native resolution, then upscale with TensorRT nodes (FL2VA + R2VA workflows).
Duration
Choose a preset: 5s / 10s / 15s. The Math Expression node snaps the frame count to the model's 17-frame-per-block grid (17k+5 at 24 fps).
๐ณ Docker / Cloud Ready
OneClick RunPod Template
Prefer a ready-to-go environment? Use the OneClick ComfyUI MiniMax H3 Qwen3VL RunPod template:
Docker image:
huchukato/comfyui-qwenvl-runpod:cu13-minimaxBase:
runpod/comfyui:cuda13.0All custom nodes pre-installed
All 4 workflows auto-downloaded at boot
Models auto-downloaded at first boot (~50 GB, persistent)
ComfyUI v0.30.0+ forced at boot
Sage Attention, FP16 accumulation, async offload
TensorRT upscaling + RIFE interpolation
ComfyUI Args (pre-configured)
--disable-auto-launch
--fast fp16_accumulation
--use-sage-attention
--reserve-vram 2
--cuda-malloc
--async-offload
๐ Why Choose ComfyUI-QwenVL-Mod + MiniMax H3?
๐ฌ For Content Creators
Native audio: Video and audio in one pass โ no separate MMAudio needed
Multilingual: Write in any language, Qwen3-VL handles translation
Professional: Official MiniMax H3 prompt format with camera vocabulary and speaker tags
Quality: 768p native, TensorRT upscale to higher resolution
๐ฅ For NSFW Content
Explicit: Uncensored generation with dedicated NSFW presets
9 presets: 3 base ๐ฌ + 3 FL2VA ๐ + 3 R2VA ๐๏ธ โ each tuned for its mode
Detailed: Rich scene descriptions with explicit diegetic soundscape
Natural: Realistic progression, consistent characters
Audio: Native moans, breaths, skin contact, ambient sounds
โก For Power Users
Customizable: Easy to modify presets and system prompts
Extendable: Add your own Qwen3-VL models (GGUF or HF)
Integrable: Works with existing ComfyUI setups
Optimized: Sage Attention, FP16, async offload, smart caching
Multi-reference:
image2input for FL2VA and R2VA workflows
๐ What Makes This Special?
First: Complete MiniMax H3 workflow pack with Qwen3-VL auto-prompting
Native audio: No separate audio node โ MiniMax H3 does it all
4 workflows: T2VA, I2VA, FL2VA, R2VA โ covers all MiniMax H3 modes
Multi-reference: Qwen3-VL sees all connected images (not just the first)
TensorRT: Built-in upscaling and frame interpolation
9 NSFW presets: Dedicated presets for each mode with correct prompt structure
Multilingual: Any input language, auto-translated and formatted
Ready: Works out-of-the-box with included workflows
๐ฏ What's New in v2.5
๐ MiniMax H3 Full Support
โ 4 workflows: T2VA, I2VA, FL2VA, R2VA โ all modes covered
โ Multi-reference input:
image2input โ Qwen3-VL sees all images as individual imagesโ 9 NSFW presets: 3 base ๐ฌ + 3 R2VA ๐๏ธ + 3 FL2VA ๐ with correct prompt structure
โ Smooth camera: All presets enforce smooth, continuous camera motion
โ Native audio: Video + stereo audio in one pass
โ Official format: 3-field (base) and 6-field (R2VA) prompt formats
๏ฟฝ Qwen3.5 Thinking Fix
โ
/no_thinkprefix for Qwen3.5 models (enable_thinking deprecated in recent llama.cpp)โ Broadened architecture detection (qwen35, qwen35moe, qwen35_vl)
โ Works across both HF and GGUF nodes
๐ฆ Workflow Organization
โ Moved workflows to
minimax/folderโ Renamed FLF to FL2VA (clearer naming)
โ Added Civitai documentation
๐ Credits
MiniMax H3 โ MiniMax ยท Comfy-Org/MiniMax-H3
ComfyUI โ comfyanonymous/ComfyUI
QwenVL-Mod โ huchukato/ComfyUI-QwenVL-Mod
Qwen3-VL โ Qwen Team / Alibaba
INT4 models โ Merserk/MiniMax-H3-INT4-ConvRot
TensorRT RIFE / Upscaler โ huchukato
VideoHelperSuite โ Kosinkadink
Easy-Use โ yolain
was-node-suite โ ltdrdata
find-perfect-resolution โ ashtar1984
๐ License
Workflows are released under the same license as the underlying models and custom nodes. See each repository for details.
MiniMax H3 model weights: Comfy-Org/MiniMax-H3 โ MiniMax H3 Community License.
Built with โค๏ธ for the ComfyUI community