โจOne-click Pod available on:โจ
๐ฃ Deploy on RunPod with CUDA 13.0

๐ฃ Deploy on RunPod with CUDA 12.8

๐ก Deploy on VastAI

๐ณ RunPod users: Just click the template link, choose a GPU, and everything installs automatically โ ComfyUI, all nodes, all workflows, and WAN 2.2 models (~30GB) download in the background on first boot. No manual setup needed. ComfyUI starts immediately while models download.
โ๏ธ buymeacoffee
IMPORTANT:
If you install RES4LYF node it will broke the MoEKSampler, to use it you have to use the KSampler included in that node.
ComfyUI-QwenVL-Mod โ Enhanced Vision-Language with WAN 2.2 Version 2.9.0 (2026/09/07) โ ๐ฌ WAN 2.2 NSFW Video + WAN Remix T2V/I2V Models + Story/Timeline Workflows (up to 20s) + SVI Camera + FL2V First-Last-Frame + 8 Workflows + Wildcards Included
โฌ๏ธ 2026/09/07 UPDATE โฌ๏ธ
๐ฆ What's Included โ 8 Workflows
All workflows are pre-wired with Qwen3-VL auto-prompting, WAN Remix diffusion models, and TensorRT upscale + RIFE interpolation where applicable.
WAN2.2-T2V-Qwen3.5.jsonโ T2V ยท Text-to-video, 5 secondsWAN2.2-I2V-Qwen3.5.jsonโ I2V ยท Image-to-video, 5 secondsWAN2.2-FL2V-Qwen3.5.jsonโ FL2V ยท First-Last-Frame to video, TensorRT upscale + RIFEWAN2.2-I2V-20s-Qwen3.5.jsonโ I2V 20s ยท Single-scene image-to-video, 20 secondsWAN2.2-I2V-20s-Story-Qwen3.5.jsonโ Story I2V ยท Multi-prompt timeline, 20 seconds (4 ร 5s)WAN2.2-I2V-SVI-20s-Qwen3.5.jsonโ SVI 20s ยท Subject Video Identity, 20 secondsWAN2.2-I2V-SVI-20s-Story-Qwen3.5.jsonโ Story SVI ยท Timeline with SVI identity lock, 20 secondsWAN2.2-T2V-I2V-Story-Qwen3.5.jsonโ Story T2V+I2V ยท Timeline mixing T2V and I2V, 20 seconds
๐ WAN Remix T2V/I2V Models โ 4 Variants
All WAN 2.2 workflows now use the WAN Remix T2V/I2V diffusion models. Download from the original Civitai pages:
wan22RemixT2VI2V_t2vHighV20.safetensorsโ T2V ยท High motion ยท ~14.3 GB โ Civitai โ WAN REMIX T2V v2.0 Highwan22RemixT2VI2V_t2vLowV20.safetensorsโ T2V ยท Low motion (stable) ยท ~14.3 GB โ Civitai โ WAN REMIX T2V v2.0 Lowwan22RemixT2VI2V_i2vHighV30.safetensorsโ I2V / FL2V / SVI / Story ยท High motion ยท ~14.3 GB โ Civitai โ WAN REMIX v2.1 FP8 Highwan22RemixT2VI2V_i2vLowV30.safetensorsโ I2V / FL2V / SVI / Story ยท Low motion (stable) ยท ~14.3 GB โ Civitai โ WAN REMIX v2.1 FP8 LowCredits: FX_FeiHou (FP8 Remix)
High vs Low: High = more dynamic camera and subject motion; Low = more stable, controlled motion (better for subtle animations)
๐งน Removed: WAN Enhanced NSFW SVI Camera
Removed
wan22EnhancedNSFWSVICamera_nsfwV2FP8H/Lmodels โ superseded by WAN RemixDocker and provisioning cleaned up
๐ฒ PMP Wildcards โ Downloaded at Boot
Wildcards (
__pmp/prmpt/*) are now downloaded from ComfyUI-Garage at boot timeNo Docker rebuild needed to update wildcards โ just push to Garage and restart the pod
comfy-tagcompleteships with wildcard fallback for local installs
โฌ๏ธ 2026/08/04 UPDATE โฌ๏ธ
โจ ComfyUI QwenVL-Mod Node Update โจ
v2.4 โ Local Model Discovery + Qwen3.5 + SageAttention
We haven't forgotten about this node! Here's what's new since v2.2:
๐ฌ LTX 2.3 Presets (v2.3): New specialized presets for LTX 2.3 I2V and T2V with official prompting guides. Multilingual support for all presets, simplified single-paragraph format (max 200 words), full NSFW support.
๐ Local Model Discovery (v2.4): Drop your GGUF/HF files into
models/LLM/and they show up in the dropdown automatically โ no more JSON editing. Auto-pairs mmproj files for vision GGUF models.๐ง Qwen3.5 support (v2.4): Architecture detection from file metadata (GGUF header / HF
config.json), automatic thinking-mode disabling, forcedtop_k=20.โก SageAttention restored (v2.4): Architecture-aware kernels (Blackwell FP8, Hopper FP8, Ada FP8, Ampere FP16) with graceful SDPA fallback.
๐ง Also updated our companion nodes:
ComfyUI-Upscaler-TensorRT-Auto โ TensorRT upscaling with auto-detection, CUDA 12/13 wheels pre-baked
ComfyUI-RIFE-TensorRT-Auto โ TensorRT frame interpolation, CUDA 12/13 wheels pre-baked
comfy-tagcomplete โ Tag completion with wildcard support for WAN 2.2 workflows
ComfyUI-HuggingFace โ Model download integration for local discovery
All WAN 2.2 and LTX 2.3 workflows (T2V, I2V, SVI, MMAudio, GGUF variants) are tested and working with the updated nodes. Grab the latest version and let us know how it goes!
โ ๏ธ Requirements โ Read First!
GPU & VRAM
๐ข Recommended โ RTX 5090 (32 GB) / RTX PRO 6000 (48 GB) / RTX 4090 (24 GB) โ FP8 Remix models
๐ก Mid-range โ RTX 3090 (24 GB) / RTX 4080 (16 GB) โ FP8 with offload
๐ Lower VRAM โ 12โ16 GB โ FP8 with aggressive offload
Model Quantization Options
FP8 (recommended) โ ~14.3 GB per diffusion model + ~4.8 GB text encoder = ~19 GB active set โ huchukato/garage
FP16 (full) โ ~42 GB per diffusion model + ~12 GB text encoder = ~54 GB total โ Comfy-Org/Wan_2.2
Software
ComfyUI: v0.31.0+
Python: 3.10+
CUDA: 12.8+ (13.0 recommended)
Storage: allow at least 80 GB for the complete provisioned package
Text Encoder
FP8 (recommended, NSFW):
nsfw_wan_umt5-xxl_fp8_scaled.safetensors(~4.8 GB) โ NSFW-API/NSFW-Wan-UMT5-XXLFP8 (standard):
umt5_xxl_fp8_e4m3fn_scaled.safetensorsโ Comfy-Org
๐ What is ComfyUI-QwenVL-Mod?
A powerful enhanced vision-language node for ComfyUI that combines Qwen3-VL models with WAN 2.2 video generation workflows. Features multilingual support, visual style detection, NSFW capabilities, Story/Timeline multi-prompt generation, and MMAudio integration.
Think: "Your all-in-one solution for intelligent prompt enhancement and video generation with WAN 2.2!"
๐ฌ Key Features
๐ WAN 2.2 Video Generation
T2V (Text-to-Video): Generate video from text prompts
I2V (Image-to-Video): Animate a first-frame image
FL2V (First-Last-Frame): Generate the transition between two keyframes โ Qwen3-VL sees both frames
SVI (Subject Video Identity): Lock character identity across generations using reference images
Story (Timeline): Multi-prompt timeline generation โ up to 4 prompts for 20-second videos with automatic scene transitions
๐ง Qwen3-VL Auto-Prompting
Multilingual: Write your prompt in any language โ Qwen3-VL translates and converts it
Auto-format: Generates optimized WAN 2.2 prompt format
Multi-reference: Qwen3-VL sees all connected images via
image+image2inputsVisual style detection: 12+ artistic styles (photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy, etc.)
Smart caching: Performance optimization with Fixed Seed Mode
GGUF backend: Efficient local model inference with quantization support
Qwen3.5 support: Thinking mode disabled via
/no_thinkfor fast prompt generationCamera tag dropdown: 19 camera movements selectable directly in the node UI
๐ต MMAudio Integration
MMAudio can be added to any workflow by connecting the MMAudio nodes to the generated video output. The node analyzes the video and produces synchronized audio (music, speech, sound effects).
๐จ NSFW Support
Comprehensive content generation without restrictions
Dedicated NSFW presets for each workflow type
Natural progression, style adaptation, consistent characters
๐ฏ QwenVL-Mod NSFW Presets
The workflows include built-in NSFW presets for the Qwen3-VL prompt enhancer:
๐ฟ T2V Presets
๐ฟ Wan 2.2 NSFW T2Vโ Standard T2V prompt๐ฟ Wan 2.2 NSFW T2V Timeline (5s)โ Timeline format for Story workflows
๐ฅ I2V Presets
๐ฅ Wan 2.2 NSFW I2V Scene (5s)โ Single scene, 5 seconds๐ Wan 2.2 NSFW I2V Scene (20s)โ Single scene, 20 seconds๐ฌ Wan 2.2 NSFW I2V Timeline (20s)โ Multi-prompt timeline, 20 seconds
๐ FL2V Presets
๐ Wan 2.2 NSFW FL2V Scene (5s)โ Transition between first and last frame
๐ผ๏ธ Utility Presets
๐ผ๏ธ Detailed Descriptionโ SFW detailed scene description (for non-NSFW use)
SFW presets are also available. Edit the preset dropdown in the QwenVL node to switch.
๐ผ๏ธ Multi-Reference Input (image2)
The QwenVL-Mod node has two image inputs:
T2V: no images needed
I2V:
image= first frameFL2V:
image= first frame,image2= last frameSVI:
image= primary reference,image2= additional references (batch)Story:
image= first frame for I2V segments,image2= optional second reference
Qwen3-VL sees all connected images as individual images, enabling proper multi-reference analysis.
๐ฎ Usage Examples
Basic Text-to-Video (T2V)
Load
WAN2.2-T2V-Qwen3.5.jsonWrite your prompt in any language
Select preset
๐ฟ Wan 2.2 NSFW T2VGenerate video
Image-to-Video (I2V)
Load
WAN2.2-I2V-Qwen3.5.jsonUpload your first-frame image to
imageSelect preset
๐ฅ Wan 2.2 NSFW I2V Scene (5s)Write what happens next (in any language)
Generate animated video
First-Last-Frame (FL2V)
Load
WAN2.2-FL2V-Qwen3.5.jsonUpload first-frame to
image, last-frame toimage2Select preset
๐ Wan 2.2 NSFW FL2V Scene (5s)Describe the transition between the two frames
Generate the interpolated video with TensorRT upscale + RIFE
Story / Timeline (I2V Story)
Load
WAN2.2-I2V-20s-Story-Qwen3.5.jsonUpload first-frame to
imageSelect preset
๐ฌ Wan 2.2 NSFW I2V Timeline (20s)Write prompts for each timeline segment (up to 4 prompts, 5s each)
Generate a 20-second video with automatic scene transitions
Recommended:
max_tokens = 2048,context_length = 16384+for 20s timelines
20-Second Single Scene (I2V 20s)
Load
WAN2.2-I2V-20s-Qwen3.5.jsonUpload first-frame to
imageSelect preset
๐ Wan 2.2 NSFW I2V Scene (20s)Write what happens next (in any language)
Generate a single-scene 20-second video
SVI โ Subject Video Identity (20s)
Load
WAN2.2-I2V-SVI-20s-Qwen3.5.jsonUpload primary reference to
image, additional references toimage2Select preset
๐ฅ Wan 2.2 NSFW I2V Scene (20s)Generate a 20-second video with locked character identity
Story SVI โ Timeline with Identity Lock (20s)
Load
WAN2.2-I2V-SVI-20s-Story-Qwen3.5.jsonUpload primary reference to
image, additional references toimage2Select preset
๏ฟฝ Wan 2.2 NSFW I2V Timeline (20s)Write prompts for each timeline segment
Generate a 20-second Story video with consistent character identity
๐ง Technical Specifications
โก Performance
Output: 720p/1080p, 16 fps (native), up to 20 seconds (Story)
Upscale: TensorRT RealESRGAN (FL2V workflow)
Frame interpolation: RIFE v4.25 โ 48 fps (FL2V workflow)
Sage Attention: FP16 accumulation, async offload
Smart caching: Reuse prompts with same inputs, Fixed Seed Mode for text-only caching
๐จ Model Support
Qwen3-VL 4B: 7 GGUF variants (2.38 GB โ 4.28 GB)
Qwen3-VL 8B: 7 GGUF variants (4.8 GB โ 8.71 GB)
Qwen3.5: 4B / 9B / 27B (uncensored, heretic, unsloth) โ thinking mode disabled
HF Models: Josiefed, official, Heretic-Stable variants
Quantization: Q4_K_S, Q5_K_S, FP16, INT8, FP8
๐ Multilingual Capabilities
Input languages: Any language supported
Auto-translation: Automatic translation to optimized English
Style detection: Works with multilingual prompts
Cultural adaptation: Context-aware prompt enhancement
๐ฆ Installation
Quick Install
Download: ComfyUI-QwenVL-Mod (latest version)
Extract to
ComfyUI/custom_nodes/ComfyUI-QwenVL-ModInstall requirements:
pip install -r requirements.txtRestart ComfyUI
Load included workflows from
wan22/folder
Custom Nodes Required
ComfyUI-QwenVL-Mod โ All workflows (Qwen3-VL prompt enhancer) โ huchukato/ComfyUI-QwenVL-Mod
ComfyUI-RIFE-TensorRT-Auto โ FL2V (frame interpolation) โ huchukato/ComfyUI-RIFE-TensorRT-Auto
ComfyUI-Upscaler-TensorRT-Auto โ FL2V (upscaling) โ huchukato/ComfyUI-Upscaler-TensorRT-Auto
ComfyUI-VideoHelperSuite โ All workflows (VHS_VideoCombine) โ Kosinkadink/ComfyUI-VideoHelperSuite
ComfyUI-Easy-Use โ FL2V (easy showAnything) โ yolain/ComfyUI-Easy-Use
ComfyUI-PerfectVideoResolution โ All workflows (resolution selector) โ huchukato/ComfyUI-PerfectVideoResolution
ComfyUI-WanMoeKSampler โ Story workflows (WanMoe advanced sampling) โ stduhpf/ComfyUI-WanMoeKSampler
ComfyUI-PainterI2V โ Story workflows (Painter I2V) โ princepainter/ComfyUI-PainterI2V
ComfyUI-PainterLongVideo โ Story workflows (Painter long video) โ princepainter/ComfyUI-PainterLongVideo
ComfyUI-mxToolkit โ Story workflows (mxSlider) โ Smirnov75/ComfyUI-mxToolkit
ComfyUI-TagComplete โ Wildcards (WildcardProcessor +
__pmp/prmpt/*) โ huchukato/comfy-tagcompletergthree-comfy โ Power Lora Loader, Fast Groups Bypasser โ rgthree/rgthree-comfy
Euler-Smea-Dyn-Sampler โ Alternative samplers โ Koishi-Star/Euler-Smea-Dyn-Sampler
Models Required
FP8 Workflows (T2V):
models/diffusion_models/โwan22RemixT2VI2V_t2vHighV20.safetensors(~14.3 GB) orwan22RemixT2VI2V_t2vLowV20.safetensorsโ huchukato/garagemodels/text_encoders/โnsfw_wan_umt5-xxl_fp8_scaled.safetensors(~4.8 GB) โ NSFW-API/NSFW-Wan-UMT5-XXLmodels/vae/โwan_2.1_vae.safetensors(~253 MB) โ Comfy-Org
FP8 Workflows (I2V / FL2V / SVI / Story):
models/diffusion_models/โwan22RemixT2VI2V_i2vHighV30.safetensors(~14.3 GB) orwan22RemixT2VI2V_i2vLowV30.safetensorsโ huchukato/garageSame text encoder + VAE as T2V
TensorRT Engines (FL2V only):
models/upscale_models/โRealESRGAN_x4(TensorRT engine)models/rife/โrife425_ensemble_False_scale_1_sim(TensorRT engine)
TensorRT engines must be built for your specific GPU. See ComfyUI-RIFE-TensorRT-Auto and ComfyUI-Upscaler-TensorRT-Auto for build instructions.
Download Links
Diffusion (FP8, T2V High): WAN REMIX T2V v2.0 High โ Civitai
Diffusion (FP8, T2V Low): WAN REMIX T2V v2.0 Low โ Civitai
Diffusion (FP8, I2V High): WAN REMIX v2.1 FP8 High โ Civitai
Diffusion (FP8, I2V Low): WAN REMIX v2.1 FP8 Low โ Civitai
Text encoder (NSFW FP8): nsfw_wan_umt5-xxl_fp8_scaled.safetensors
Text encoder (standard FP8): umt5_xxl_fp8_e4m3fn_scaled.safetensors
๐ฌ WAN 2.2 Prompting Notes
How to Write Your Prompt
Describe the scene naturally. Be clear about the concepts below โ Qwen3-VL handles the rest:
๐จ Visual style (put it first):
photorealistic,cinematic,anime,3D CG,claymation,vintage film,watercolor,fantasy๐ฅ Subjects: number, gender, appearance, clothing, position, expression
๐ Action / motion: what happens, speed, interaction
๐ฅ Camera: dolly, pan, zoom, static, handheld, crane, orbit โ smooth and continuous
๐ Environment: setting, lighting, atmosphere, time of day
๐ Audio (optional): connect MMAudio nodes to add synchronized sound
๐ FL2V: Describe the transition between frames, not the scene (images fix the scene) ๐ Story: Write separate prompts for each timeline segment โ Qwen3-VL handles the transitions
Resolution Guidance
WAN 2.2 native resolutions:
๐ฑ Portrait: 832ร1216 ยท 720ร1280
โฌ Square: 1024ร1024
๐ฅ๏ธ Landscape: 1216ร832 ยท 1280ร720
โ ๏ธ Match the aspect ratio to your input image! Forcing 16:9 on a portrait image will squash it.
Duration
Standard: 5 seconds (81 frames at 16 fps)
Story/Timeline: up to 20 seconds (4 ร 5s segments)
Frame interpolation: RIFE doubles framerate to 48 fps where applicable
๐ฅ Camera Control Tags
All WAN 2.2 NSFW presets support camera control via the camera_tag dropdown on the QwenVL node โ no need to type tags manually. Select from 19 camera movements:
[STATIC_CAMERA]/[LOCKED_OFF]โ Camera completely static[SLOW_ZOOM_IN]โ Slow continuous push-in[SLOW_ZOOM_OUT]โ Slow continuous pull-back[FAST_ZOOM_IN]โ Fast aggressive push-in[FAST_ZOOM_OUT]โ Fast pull-back, reveal context[PAN_LEFT]/[PAN_RIGHT]โ Smooth horizontal pan[TILT_UP]/[TILT_DOWN]โ Smooth vertical tilt[DOLLY_IN]/[DOLLY_OUT]โ Physical dolly movement (parallax)[TRACKING_LEFT]/[TRACKING_RIGHT]โ Lateral tracking shot[CRANE_UP]/[CRANE_DOWN]โ Crane/jib movement[ORBIT]โ Smooth 360-degree orbit around subject[HANDHELD]โ Subtle handheld sway with micro-movements[ROLL]โ Slow camera roll (rotation around lens axis)
How it works: the selected tag is injected at the start of the prompt AND as a FINAL CAMERA DIRECTIVE at the end, so Qwen respects it despite recency bias. The subject stays alive and active โ the tag controls only the camera.
๐ฒ Wildcards
Selected workflows include a WildcardProcessor node that injects randomized prompt fragments from the PMP's Prompt Engine (__pmp/prmpt/*) wildcard library.
How It Works
The WildcardProcessor node sits before the Qwen3-VL prompt enhancer
At queue time, each
__wildcard__token is replaced with a random line from the corresponding.txtfileThe expanded text is passed to Qwen3-VL, which converts it into the WAN 2.2 prompt format
Different seed = different wildcard picks โ use a fixed seed for reproducible results
Customizing Wildcards
Edit existing: open the
.txtfiles underComfyUI/custom_nodes/comfy-tagcomplete/wildcards/pmp/prmpt/Add your own: create a new
.txtfile, e.g.pmp/prmpt/mytags.txt, then reference it as__pmp/prmpt/mytags__Remove a wildcard: delete the
__...__token from the WildcardProcessortextfieldDisable randomization: replace the
__wildcard__token with a fixed string
Required Custom Node
ComfyUI-TagComplete (includes the
WildcardProcessornode and the__pmp/prmpt/*wildcard set) โ huchukato/comfy-tagcomplete
The wildcard files ship with the custom node as fallback. On Docker/Vast.ai deployments, wildcards are downloaded from ComfyUI-Garage at boot for the latest version.
๐ณ Docker / Cloud Ready
OneClick RunPod Template
Prefer a ready-to-go environment? Use the OneClick - ComfyUI - WAN 2.2 - Qwen3VL RunPod template:
Docker image:
huchukato/comfyui-qwenvl-runpod:cu13-wan22(CUDA 13.0) orhuchukato/comfyui-qwenvl-runpod:cu128-wan22(CUDA 12.8)Base:
huchukato/comfyui-base:cu130All custom nodes pre-installed
ComfyUI Args:
--disable-auto-launch --fast fp16_accumulation --use-sage-attention --cuda-malloc --async-offloadAll 8 workflows auto-downloaded at boot
Models auto-downloaded at first boot (~62 GB including 4 WAN Remix diffusion models, NSFW text encoder, VAE; persistent)
ComfyUI v0.34.2 baked into base image
Sage Attention, FP16 accumulation, async offload
TensorRT upscaling + RIFE interpolation
PMP wildcards auto-downloaded from Garage at boot
Access: ComfyUI
:8188ยท JupyterLab:8888ยท FileBrowser:8080(useradmin/ passwordadminadmin12) ยท SSHssh root@pod-ip
Vast.ai Provisioning
A Vast.ai provisioning script is also available:
Script:
vastai/wan22-provisioning.shDownloads all models, workflows, wildcards, and custom nodes on first boot
Same model set as RunPod Docker
ComfyUI Args (pre-configured)
--disable-auto-launch
--fast fp16_accumulation
--use-sage-attention
--cuda-malloc
--async-offload
๐ Why Choose ComfyUI-QwenVL-Mod + WAN 2.2?
๐ฌ For Content Creators
Multilingual: Write in any language, Qwen3-VL handles translation
Story/Timeline: Multi-prompt timelines for long-form content (up to 20s)
Quality: Native resolution, TensorRT upscale to higher resolution
๐ฅ For NSFW Content
Explicit: Uncensored generation with dedicated NSFW presets
Multiple presets: T2V, I2V (5s/20s), FL2V, Timeline โ each tuned for its mode
Detailed: Rich scene descriptions with explicit action
Natural: Realistic progression, consistent characters
โก For Power Users
Customizable: Easy to modify presets and system prompts
Extendable: Add your own Qwen3-VL models (GGUF or HF)
Optimized: Sage Attention, FP16, async offload, smart caching
Multi-reference:
image2input for FL2V and SVI workflowsStory: WanMoeKSampler + PainterI2V for complex multi-scene generation
๐ What Makes This Special?
Complete: 8 workflows covering T2V, I2V, FL2V, SVI, and Story
Auto-prompting: Qwen3-VL handles prompt enhancement in any language
Timeline: Multi-prompt Story workflows for up to 20-second videos
TensorRT: Built-in upscaling and frame interpolation
NSFW presets: Dedicated presets for each workflow type
Wildcards: PMP prompt engine for randomized variation
Docker-ready: OneClick RunPod template + Vast.ai provisioning
๐ Credits
WAN 2.2 โ Wan-AI ยท Comfy-Org/Wan_2.2
ComfyUI โ comfyanonymous/ComfyUI
QwenVL-Mod โ huchukato/ComfyUI-QwenVL-Mod
Qwen3-VL โ Qwen Team / Alibaba
WAN Remix FP8 โ FX_FeiHou
NSFW Text Encoder โ NSFW-API/NSFW-Wan-UMT5-XXL
TensorRT RIFE / Upscaler โ huchukato
VideoHelperSuite โ Kosinkadink
Easy-Use โ yolain
PerfectVideoResolution โ huchukato
WanMoeKSampler โ stduhpf
PainterI2V / PainterLongVideo โ princepainter
mxToolkit โ Smirnov75
rgthree-comfy โ rgthree
๐ License
Workflows are released under the same license as the underlying models and custom nodes. See each repository for details.
WAN 2.2 model weights: Wan-AI โ Apache 2.0.
Built with โค๏ธ for the ComfyUI community
Description
Removed the Tensorrt Upscaler and replaced with 2xLexicaRRDBNet
FAQ
Comments (41)
AILab_QwenVL_GGUF_Advanced
[QwenVL] llama_cpp is not available. Install the GGUF vision dependency first. See docs/GGUF_MANUAL_INSTALL.md
I wrote the guide to install it here and also in the WF...
@huchukatoย Yeah,the guide isn't geared toward everyone. You might have come up with something cool here, but the explanations on how to get llama going are extremely lacking.
@meritrash6350ย I wrote all the steps ypu have to do to install llama, the only thing that lack is how to activate the virtual enviroment on Windows coz I don't use Windows, I'm on Mac and Linux
@huchukatoย So, you get my point then?
@meritrash6350ย I got it but you don't have to install the GGUF version of the WF in every cases, go with the normal one where you don't have to install llama python, I cannot tell you how to get in a venv in Windows not having Windows on my PC :\
@meritrash6350 In your ComfyUI root installation folder activate enviornment using either script:
Command Prompt: \venv\Scripts\activate.bat
PowerShell: \venv\Scripts\Activate.ps1
@evantopsmithย Sorry, wasn't ignoring you, I got too busy. I'll take a look at it, and I appreciate the extra effort. I personally find the effort of dealing with venv to be the thing I hate most about Python. That and the fact it insists on caching on the C: even though that's not where I told it to install.
@evantopsmithย Thanks a lot
I am so confused on what you mean to install. The guide isn't clear, whad you you even mean by start a ComfyUI virtual environment???
When you install ComfyUI the installer create a virtual enviroment to run it, usually thers is a "venv" folder inside ComfyUI. This allows Comfy to run on a specific Python version and to install all the dependecies that it needs not for all you system but just for Comfy, to avoid comflicts in you system
Hello, in my QWENVL node, the generated prompt are completely unrelated to the prompt I entered. The generated prompt seem to only describe the content of the image itself. Why is this happening? Full i2v longvideo gguf
I am having the same issue
I am about to lose my fucking mind with this Qwen3 VL GGUF auto prompt bullshit. git cloned QwenVL-Mod (no QwenVL node from manager) start ComfyUI install dependencies, quit. download llama_cpp_python-0.3.23+cu128.basic-cp312-win_amd64.whl (Python 3.12.10, CUDA 2.8.0cu128) cmd in python_embedded (equivalent to being in active virtual environment using comfyui-easy-install) pip install --upgrade --force-reinstall llama_cpp_python-0.3.23+cu128.basic-cp312-win_amd64.whl (whl file is in same folder cmd was started in) restart comfyui and try to use Huihui-Qwen3-VL-4B-Instruct-abliterated-Q8_0.gguf.....
AILab_QwenVL_GGUF_Advanced
[QwenVL] Missing Qwen VL chat handler in llama_cpp. Install the correct fork/wheel. See docs/GGUF_MANUAL_INSTALL.md
:\
https://github.com/1038lab/ComfyUI-QwenVL
This one worked for me
@gregariousbuttons257ย This is the original node but you have to manual install the abliterated Qwen models, the only difference with mine is that I added the uncensored models
@huchukatoย but I followed all of that to the tee. which part of what I described am I doing wrong? do I have the wrong whl for my python cuda versions?
Was so hyped for this workflow but I have been trying to get it to work for the last 3 hours. I have yet to make even one generation! something with the Ksampers? Which I don't understand because they work normally in other workflows but for some reason this one just refuses to work for me!
this is what i keep getting:
Failed to validate prompt for output 1431:
* KSamplerAdvanced 1252:1269:
- Return type mismatch between linked nodes: scheduler, received_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent']) mismatch input_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent', 'beta57'])
* KSamplerAdvanced 1252:1270:
- Return type mismatch between linked nodes: scheduler, received_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent']) mismatch input_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent', 'beta57'])
Output will be ignored
Failed to validate prompt for output 1252:1259:
Output will be ignored
Failed to validate prompt for output 1327:
Output will be ignored
[QwenVL] Flash-Attn auto mode: dependency not ready, using SDPA
Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592
Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592
Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592
Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592
Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592
Processing interrupted
Prompt executed in 524.73 seconds
Seems there is a problem with the setnode with the scheduler, which version are you using?
@huchukato, i have the same problem, what do you mean by Version? Comfy?ย
@jarigoni949ย Comfy and also the WF version, coz for me all versions works on Comfy 0.10.0, Python 3.12.12 and CU13.0 with Pytorch 2.9.1
@huchukatoย Hey :)!
I'am on: Comfy 0.11.0 / Python 3.12.10 / pytorch version: 2.9.1+cu130 /
@jarigoni949ย I'm testing the WF on 0.11.0 now and works :\ You can try to do a thing: disconnect the scheduler and sampler from the subgraph and manually set them inside it, maybe it's a problem of the Selectors node
I agree. If it's not about sampler issue, it's about QWEN abridged. The model doesn't download automatically, and when you do manually, the node won't detect it and download other models.
im using comfy11.1, Python 3.12.11, 2.10.0+cu130 and CUDA13.0
@yamyproย The WF is tested on Python 3.12.12, update your py version maybe
@huchukatoย i know this is a dumb question but how do i upgrade it?
I still have the error: The checkpoint you are trying to load has model type qwen3_vl but Transformers does not recognize this architecture. This could be because of an issue with the checkpoint, or because your version of Transformers is out of date.
You updated my Qwen node? Now it supports both transformers>5.0 and <5.0, I fixed the deprecated syntax
Ive got it working with no major issues, the only problem that i have is that the generated prompts are very censored, it doesnt describe any nsfw behaviors and at most it implies vaguely that some "intimacy" might be going on. Any idea on how to make it as detailed and nsfw as possible?
If you are using my modified node and one of the Qwen Abliterated models it should work, for me it works
[QwenVL] Failed to apply SageAttention patch: module 'transformers.models.qwen2.modeling_qwen2' has no attribute 'F'
[QwenVL] SageAttention patch failed, continuing with SDPA
Is this a SageAttention // Transformers version issue?
you have sageattention==2.2.0? If it fail to load it it uses SDPA by default
Name: sageattention
Version: 2.2.0+cu130torch2.9.0andhigher.post4
I do yeah, and SDPA still only takes a moment to run QwenVL, but it's still bothering me.
I've got things tuned so that this workflow is working really well for me. My main issue now is that UpscalerTensorRT at 2x upscale on a 10 second clip is using an insane amount of vram and spilling into shared memory causing it to take upwards of 30min to run, when the rest of the generation only takes ~10 minutes. (RTX 4090) Trying to figure out if it's possible to optimize UpscalerTensorRT's memory usage or if some models loaded that don't need to be when the workflow gets there.
Edit: I moved the post-processing (UpscalerTensorRT and Interpolation) into a separate workflow and it finishes in around a minute on my video output, so there's something that I don't have the understanding to fix regarding memory management to run your entire workflow efficiently.
Oh, and thanks for your responses and thanks very much for sharing the workflow!!!
@dukefan6842872ย no problem, btw now I'm getting the same error LOL, I will take a look tomorrow. Regarding Tensorrt, try the normal Upscale and check if it takes less time. PS. I just update again the node, now have more aderence in NSFW prompting, in minutes I also release the SVI version of the WF
@dukefan6842872ย I think I fixed the problem, update the node and let me know thank u