๐ฃ Deploy on Runpod
๐ก Deploy on Vast.ai

ComfyUI-QwenVL-Mod โ Enhanced Vision-Language with LTX 2.5 Version 2.5.1 (2026/08/25) โ ๐ฌ LTX 2.5 Native Video+Audio + DeepNeuralNerd Uncensored Gemma 4 + Qwen3.5 Auto-Prompting
โ ๏ธ Requirements โ Read First!
GPU & VRAM
๐ข Recommended โ RTX 5090 / 4090 (24 GB+) โ INT8 ConvRot transformer + text encoder โ Fast, best quality
๐ก Enthusiast โ RTX 3090 / 4080 (16-24 GB) โ offload text encoder / VAE โ Slower, good quality
๐ Budget โ RTX 4060 Ti 16 GB โ aggressive offload + lower resolution โ Experimental, slow
12 GB GPUs: Not recommended for LTX 2.5 with the full DeepNeuralNerd setup. The text encoder alone is ~13.2 GB INT8.
Software
ComfyUI: v0.34.2 (required for LTX 2.5 native audio, UNETLoader, and Sage Attention flags)
Python: 3.10+
CUDA: 13.0 recommended
Storage: 120 GB+ SSD (models ~50 GB + outputs)
Qwen3-VL Prompt Enhancer
GGUF: Q4_K_S (~4.8 GB) or Q5_K_S (~5.5 GB) for 8B model
HF: Qwen3-VL-8B-Heretic-Stable (~16 GB) or Qwen3-VL-4B (~8 GB)
๐ What is ComfyUI-QwenVL-Mod?
A powerful enhanced vision-language node for ComfyUI that combines Qwen3-VL models with LTX 2.5 video generation workflows. Features multilingual support, visual style detection, native synchronized audio, and NSFW capabilities for professional AI content creation.
Think: "Your all-in-one solution for intelligent prompt enhancement and uncensored video+audio generation with LTX 2.5!"
๐ฌ Key Features
๐ LTX 2.5 Native Video+Audio Generation
I2VA (Image-to-Video+Audio): Animate a first-frame image with synchronized audio
FL2VA (First-Last-Frame): Generate the transition between two keyframes with audio โ Qwen3-VL sees both frames
LTX 2.5 produces video and audio in the same forward pass: no separate MMAudio or audio-VAE stitching is required.
๐ง Qwen3.5 Auto-Prompting
Multilingual: Write your prompt in any language โ Qwen3-VL translates and converts it
Auto-format: Generates optimized LTX 2.5 prompt structure
Multi-reference: Qwen3-VL sees all connected images via
image+image2inputsVisual style detection: 12+ artistic styles (photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy, etc.)
Smart caching: Performance optimization with Fixed Seed Mode
GGUF backend: Efficient local model inference with quantization support
Qwen3.5 support: Thinking mode disabled via
/no_thinkfor fast prompt generation
๐ Native Synchronized Audio
No separate audio node needed โ LTX 2.5 generates video and audio jointly
Voice, sound effects, ambience and music modeled together
Describe sounds and dialogue in your prompt and the model generates them natively
๐จ NSFW Support
Comprehensive content generation without restrictions
Dedicated NSFW presets tuned for LTX 2.5 I2V and FL2V
Natural progression, consistent characters, explicit diegetic soundscape
๐ฆ What's Included โ 2 Workflows
๐ผ๏ธ I2VA/T2VA โ
LTX25-I2VA-T2VA-Qwen3.5.jsonโ text + optional first-frame image (image) โ Image-to-Video+Audio or Text-to-Video+Audio. A boolean toggle switches between I2VA (with image input) and T2VA (text-only). DeepNeuralNerd uncensored Gemma 4 text encoder + Gemma 4 E2B prompt enhancer, dual-stage with latent spatial upscaler.๐ FL2VA โ
LTX25-FL2VA-Qwen3.5.jsonโ text + first-frame (image) + last-frame (image2) โ First-Last-Frame to video with audio. Qwen3-VL sees both frames and describes the transition. Includes TensorRT upscale + RIFE frame interpolation for high-resolution output.
I2VA and FL2VA include TensorRT upscaling (RealESRGAN x4) and RIFE frame interpolation.
๐ผ๏ธ Multi-Reference Input (image2)
The QwenVL-Mod node has two image inputs:
I2VA:
image= first frameFL2VA:
image= first frame,image2= last frame,frame_count = 1
Qwen3-VL sees all connected images as individual images (not as a video sequence), enabling proper multi-reference analysis for FL2VA.
๐ฏ QwenVL-Mod NSFW Presets
The workflows include built-in NSFW presets for the Qwen3-VL prompt enhancer:
๐ฌ LTX 2.5 NSFW I2Vโ Image-to-video with mandatory audio instructions๐ LTX 2.5 NSFW FL2Vโ First-last-frame transition with audio instructions
What the presets produce
๐ฌ I2VA: Rich scene description, camera motion, subject action, explicit diegetic soundscape, non-diegetic music flag
๐ FL2VA: Describes the transition path between first and last frames (not the scene โ images fix the scene). Favors single continuous shot.
All presets: smooth, continuous camera motion (no abrupt or stepped changes), explicit audio notes
SFW presets are also available. Edit the preset dropdown in the QwenVL node to switch.
๐ฎ Usage Examples
Image-to-Video (I2VA)
Load
LTX25-I2VA-Qwen3.5.jsonUpload your first-frame image to
imageSelect preset
๐ฌ LTX 2.5 NSFW I2VWrite what happens next (in any language)
Generate animated video with synchronized audio
First-Last-Frame (FL2VA)
Load
LTX25-FL2VA-Qwen3.5.jsonUpload first-frame to
image, last-frame toimage2, setframe_count=1Select preset
๐ LTX 2.5 NSFW FL2VDescribe the transition between the two frames
Generate the interpolated video with TensorRT upscale + RIFE
๐ง Technical Specifications
โก Performance
Output: Native LTX 2.5 resolution, 24 fps
Audio: Native synchronized stereo audio
Upscale: TensorRT RealESRGAN x4 (both workflows) + latent spatial upscaler x2 (I2VA)
Frame interpolation: RIFE rife49 โ higher fps (both workflows)
Sage Attention: FP16 accumulation, async offload
Smart caching: Reuse prompts with same inputs, Fixed Seed Mode for text-only caching
๐จ Model Support
Qwen3-VL 4B: 7 GGUF variants (2.38 GB โ 4.28 GB)
Qwen3-VL 8B: 7 GGUF variants (4.8 GB โ 8.71 GB)
Qwen3.5: 4B / 9B / 27B (uncensored, heretic, unsloth) โ thinking mode disabled
HF Models: Josiefed, official, Heretic-Stable variants
Quantization: Q4_K_S, Q5_K_S, FP16, INT8
๐ Multilingual Capabilities
Input languages: Any language supported
Auto-translation: Automatic translation to optimized English
Style detection: Works with multilingual prompts
Cultural adaptation: Context-aware prompt enhancement
๐ฆ Installation
Quick Install
Download: ComfyUI-QwenVL-Mod (latest version)
Extract to
ComfyUI/custom_nodes/ComfyUI-QwenVL-ModInstall requirements:
pip install -r requirements.txtRestart ComfyUI
Load included workflows from
ltx/25/folder
Custom Nodes Required
ComfyUI-QwenVL-Mod โ All workflows (Qwen3-VL prompt enhancer) โ huchukato/ComfyUI-QwenVL-Mod
ComfyUI-LTXVideo โ LTX 2.5 native nodes โ Lightricks/ComfyUI-LTXVideo
ComfyUI-RIFE-TensorRT-Auto โ FL2VA, I2VA (frame interpolation) โ huchukato/ComfyUI-RIFE-TensorRT-Auto
ComfyUI-Upscaler-TensorRT-Auto โ FL2VA, I2VA (upscaling) โ huchukato/ComfyUI-Upscaler-TensorRT-Auto
ComfyUI-VideoHelperSuite โ (VHS_VideoCombine) โ Kosinkadink/ComfyUI-VideoHelperSuite
ComfyUI-Easy-Use โ (easy showAnything) โ yolain/ComfyUI-Easy-Use
comfyui-find-perfect-resolution โ All workflows (ResolutionSelector) โ ashtar1984/comfyui-find-perfect-resolution
was-node-suite-comfyui โ FL2VA, I2VA (GetImageSize, ResizeImageMaskNode) โ ltdrdata/was-node-suite-comfyui
Note:
ComfyMathExpressionis built into ComfyUI core (v0.24.1+) โ no custom node needed.
Models Required
LTX 2.5 I2VA / FL2VA use the DeepNeuralNerd uncensored Gemma 4 setup:
models/diffusion_models/โltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors(~14 GB) โ Lightricks/LTX-2.5models/text_encoders/โgemma4-12b-uncensored-heretic-ltx2.5-comfy-int8-convrot.safetensors(~13.2 GB) โ DeepNeuralNerd/Gemma-4-12B-it-uncensored-heretic-DeepNeuralNerd-LTX_2.5_ComfyUImodels/text_encoders/โgemma4_e2b_it_bf16.safetensors(~5 GB, I2VA prompt enhancer only) โ TrevorJS/gemma-4-E2B-it-uncensoredmodels/vae/โltx-2.5-video-vae-bf16.safetensors(~500 MB) โ Lightricks/LTX-2.5models/vae/โltx-2.5-audio-vae-bf16.safetensors(~500 MB) โ Lightricks/LTX-2.5models/latent_upscale_models/โltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors(~990 MB, I2VA only) โ Lightricks/LTX-2.5
Qwen3-VL Prompt Enhancer
models/LLM/โQwen3-VL-8B-Heretic-Stable(GGUF or HF)
TensorRT Engines
models/upscale_models/โRealESRGAN_x4(TensorRT engine)models/rife/โrife49_ensemble_True_scale_1_sim(TensorRT engine)
TensorRT engines must be built for your specific GPU. See ComfyUI-RIFE-TensorRT-Auto and ComfyUI-Upscaler-TensorRT-Auto for build instructions.
๐ณ Docker / Cloud Ready
OneClick RunPod Template
Prefer a ready-to-go environment? Use the OneClick - ComfyUI - LTX 2.5 Uncensored - CU13 RunPod template:
Docker image:
huchukato/comfyui-qwenvl-runpod:cu13-ltx25Base:
huchukato/comfyui-base:cu130All custom nodes pre-installed
Both workflows auto-downloaded at boot
Models auto-downloaded at first boot (~50 GB, persistent)
ComfyUI v0.34.2 baked into base image
Sage Attention, FP16 accumulation, async offload
TensorRT upscaling + RIFE frame interpolation
ComfyUI Args (pre-configured)
--disable-auto-launch
--fast fp16_accumulation
--use-sage-attention
--cuda-malloc
--async-offload
๐ Why Choose ComfyUI-QwenVL-Mod + LTX 2.5?
๐ฌ For Content Creators
Native audio: Video and audio in one pass โ no separate audio node
Multilingual: Write in any language, Qwen3-VL handles translation
Professional: Smooth camera vocabulary and natural motion
Quality: DeepNeuralNerd uncensored Gemma 4 + TensorRT upscale
๐ฅ For NSFW Content
Explicit: Uncensored generation with dedicated NSFW presets
Detailed: Rich scene descriptions with explicit diegetic soundscape
Natural: Realistic progression, consistent characters
Audio: Native moans, breaths, skin contact, ambient sounds
โก For Power Users
Customizable: Easy to modify presets and system prompts
Extendable: Add your own Qwen3-VL models (GGUF or HF)
Integrable: Works with existing ComfyUI setups
Optimized: Sage Attention, FP16, async offload, smart caching
Multi-reference:
image2input for FL2VA workflow
๐ What Makes This Special?
First: Complete LTX 2.5 uncensored workflow pack with Qwen3-VL auto-prompting
Native audio: No separate audio node โ LTX 2.5 does it all
2 workflows: I2VA + FL2VA covering the main LTX 2.5 use cases
Multi-reference: Qwen3-VL sees all connected images (not just the first)
TensorRT: Built-in upscaling and frame interpolation
NSFW presets: Dedicated presets tuned for LTX 2.5
Multilingual: Any input language, auto-translated and formatted
Ready: Works out-of-the-box with included workflows
๐ฏ What's New in v2.5.1
โ LTX 2.5 DeepNeuralNerd Uncensored Setup
Switched to DeepNeuralNerd uncensored heretic Gemma 4 12B text encoder (INT8 ConvRot)
Uses LTX 2.5 22B distilled transformer (INT8 ConvRot) from Lightricks
Separate video VAE and audio VAE loaders (LTX 2.5 architecture)
Added Gemma 4 E2B uncensored prompt enhancer (TrevorJS) for I2VA
Added latent spatial upscaler x2 for I2VA dual-stage upscaling
โ Architecture Changes from LTX 2.3
LTX 2.5 uses
UNETLoader+VAELoader(transformer and VAE are separate files)No CheckpointLoaderSimple โ VAE is not embedded in the checkpoint
No DMD LoRA โ LTX 2.5 distilled transformer has built-in distillation
โ Native Audio Workflows
Both I2VA and FL2VA workflows generate synchronized native audio
Presets include mandatory audio instructions
โ Qwen3.5 Rename
Workflow filenames updated from
Qwen3VLtoQwen3.5for consistency
๐ Credits
ComfyUI โ comfyanonymous/ComfyUI
LTX 2.5 โ Lightricks/LTX-2.5
QwenVL-Mod โ huchukato/ComfyUI-QwenVL-Mod
Qwen3-VL โ Qwen Team / Alibaba
DeepNeuralNerd Uncensored Gemma 4 โ DeepNeuralNerd
Gemma 4 E2B Uncensored โ TrevorJS
LTXVideo nodes โ Lightricks/ComfyUI-LTXVideo
TensorRT RIFE / Upscaler โ huchukato
VideoHelperSuite โ Kosinkadink
Easy-Use โ yolain
find-perfect-resolution โ ashtar1984
was-node-suite โ ltdrdata
๐ License
Workflows are released under the same license as the underlying models and custom nodes. See each repository for details.
LTX 2.5 model weights: Lightricks/LTX-2.5 โ Lightricks Community License.
Gemma 4 model weights: DeepNeuralNerd โ see repository for license details.
Built with โค๏ธ for the ComfyUI community
Description
FAQ
Comments (2)
Hi there, in the workflow the prompt doesnt have the LTX2.3 or LTX2.5 preset prompts
๐ฅ LTX 2.3 NSFW I2V (5s, 10s, 20s)
๐ LTX 2.3 NSFW FL2VA (5s, 10s, 20s)
๐ฅ LTX 2.3 NSFW T2V (5s, 10s)
these are the presets, works also on 2.5