✨One-click Pod available on:✨
🟣 Deploy on RunPod with CUDA 13.0

🟣 Deploy on RunPod with CUDA 12.8

🟡 Deploy on VastAI

🐳 RunPod users: Just click the template link, choose a GPU, and everything installs automatically — ComfyUI, all nodes, all workflows, and WAN 2.2 models (~30GB) download in the background on first boot. No manual setup needed. ComfyUI starts immediately while models download.
☕️ buymeacoffee
IMPORTANT:
If you install RES4LYF node it will broke the MoEKSampler, to use it you have to use the KSampler included in that node.
ComfyUI-QwenVL-Mod — Enhanced Vision-Language with WAN 2.2 Version 2.9.0 (2026/09/07) — 🎬 WAN 2.2 NSFW Video + WAN Remix T2V/I2V Models + Story/Timeline Workflows (up to 20s) + SVI Camera + FL2V First-Last-Frame + 8 Workflows + Wildcards Included
⬆️ 2026/09/07 UPDATE ⬆️
📦 What's Included — 8 Workflows
All workflows are pre-wired with Qwen3-VL auto-prompting, WAN Remix diffusion models, and TensorRT upscale + RIFE interpolation where applicable.
WAN2.2-T2V-Qwen3.5.json— T2V · Text-to-video, 5 secondsWAN2.2-I2V-Qwen3.5.json— I2V · Image-to-video, 5 secondsWAN2.2-FL2V-Qwen3.5.json— FL2V · First-Last-Frame to video, TensorRT upscale + RIFEWAN2.2-I2V-20s-Qwen3.5.json— I2V 20s · Single-scene image-to-video, 20 secondsWAN2.2-I2V-20s-Story-Qwen3.5.json— Story I2V · Multi-prompt timeline, 20 seconds (4 × 5s)WAN2.2-I2V-SVI-20s-Qwen3.5.json— SVI 20s · Subject Video Identity, 20 secondsWAN2.2-I2V-SVI-20s-Story-Qwen3.5.json— Story SVI · Timeline with SVI identity lock, 20 secondsWAN2.2-T2V-I2V-Story-Qwen3.5.json— Story T2V+I2V · Timeline mixing T2V and I2V, 20 seconds
🔄 WAN Remix T2V/I2V Models — 4 Variants
All WAN 2.2 workflows now use the WAN Remix T2V/I2V diffusion models. Download from the original Civitai pages:
wan22RemixT2VI2V_t2vHighV20.safetensors— T2V · High motion · ~14.3 GB — Civitai — WAN REMIX T2V v2.0 Highwan22RemixT2VI2V_t2vLowV20.safetensors— T2V · Low motion (stable) · ~14.3 GB — Civitai — WAN REMIX T2V v2.0 Lowwan22RemixT2VI2V_i2vHighV30.safetensors— I2V / FL2V / SVI / Story · High motion · ~14.3 GB — Civitai — WAN REMIX v2.1 FP8 Highwan22RemixT2VI2V_i2vLowV30.safetensors— I2V / FL2V / SVI / Story · Low motion (stable) · ~14.3 GB — Civitai — WAN REMIX v2.1 FP8 LowCredits: FX_FeiHou (FP8 Remix)
High vs Low: High = more dynamic camera and subject motion; Low = more stable, controlled motion (better for subtle animations)
🧹 Removed: WAN Enhanced NSFW SVI Camera
Removed
wan22EnhancedNSFWSVICamera_nsfwV2FP8H/Lmodels — superseded by WAN RemixDocker and provisioning cleaned up
🎲 PMP Wildcards — Downloaded at Boot
Wildcards (
__pmp/prmpt/*) are now downloaded from ComfyUI-Garage at boot timeNo Docker rebuild needed to update wildcards — just push to Garage and restart the pod
comfy-tagcompleteships with wildcard fallback for local installs
⬆️ 2026/08/04 UPDATE ⬆️
✨ ComfyUI QwenVL-Mod Node Update ✨
v2.4 — Local Model Discovery + Qwen3.5 + SageAttention
We haven't forgotten about this node! Here's what's new since v2.2:
🎬 LTX 2.3 Presets (v2.3): New specialized presets for LTX 2.3 I2V and T2V with official prompting guides. Multilingual support for all presets, simplified single-paragraph format (max 200 words), full NSFW support.
🔍 Local Model Discovery (v2.4): Drop your GGUF/HF files into
models/LLM/and they show up in the dropdown automatically — no more JSON editing. Auto-pairs mmproj files for vision GGUF models.🧠 Qwen3.5 support (v2.4): Architecture detection from file metadata (GGUF header / HF
config.json), automatic thinking-mode disabling, forcedtop_k=20.⚡ SageAttention restored (v2.4): Architecture-aware kernels (Blackwell FP8, Hopper FP8, Ada FP8, Ampere FP16) with graceful SDPA fallback.
🔧 Also updated our companion nodes:
ComfyUI-Upscaler-TensorRT-Auto — TensorRT upscaling with auto-detection, CUDA 12/13 wheels pre-baked
ComfyUI-RIFE-TensorRT-Auto — TensorRT frame interpolation, CUDA 12/13 wheels pre-baked
comfy-tagcomplete — Tag completion with wildcard support for WAN 2.2 workflows
ComfyUI-HuggingFace — Model download integration for local discovery
All WAN 2.2 and LTX 2.3 workflows (T2V, I2V, SVI, MMAudio, GGUF variants) are tested and working with the updated nodes. Grab the latest version and let us know how it goes!
⚠️ Requirements — Read First!
GPU & VRAM
🟢 Recommended — RTX 5090 (32 GB) / RTX PRO 6000 (48 GB) / RTX 4090 (24 GB) → FP8 Remix models
🟡 Mid-range — RTX 3090 (24 GB) / RTX 4080 (16 GB) → FP8 with offload
🟠 Lower VRAM — 12–16 GB → FP8 with aggressive offload
Model Quantization Options
FP8 (recommended) — ~14.3 GB per diffusion model + ~4.8 GB text encoder = ~19 GB active set → huchukato/garage
FP16 (full) — ~42 GB per diffusion model + ~12 GB text encoder = ~54 GB total → Comfy-Org/Wan_2.2
Software
ComfyUI: v0.31.0+
Python: 3.10+
CUDA: 12.8+ (13.0 recommended)
Storage: allow at least 80 GB for the complete provisioned package
Text Encoder
FP8 (recommended, NSFW):
nsfw_wan_umt5-xxl_fp8_scaled.safetensors(~4.8 GB) — NSFW-API/NSFW-Wan-UMT5-XXLFP8 (standard):
umt5_xxl_fp8_e4m3fn_scaled.safetensors— Comfy-Org
🌟 What is ComfyUI-QwenVL-Mod?
A powerful enhanced vision-language node for ComfyUI that combines Qwen3-VL models with WAN 2.2 video generation workflows. Features multilingual support, visual style detection, NSFW capabilities, Story/Timeline multi-prompt generation, and MMAudio integration.
Think: "Your all-in-one solution for intelligent prompt enhancement and video generation with WAN 2.2!"
🎬 Key Features
🚀 WAN 2.2 Video Generation
T2V (Text-to-Video): Generate video from text prompts
I2V (Image-to-Video): Animate a first-frame image
FL2V (First-Last-Frame): Generate the transition between two keyframes — Qwen3-VL sees both frames
SVI (Subject Video Identity): Lock character identity across generations using reference images
Story (Timeline): Multi-prompt timeline generation — up to 4 prompts for 20-second videos with automatic scene transitions
🧠 Qwen3-VL Auto-Prompting
Multilingual: Write your prompt in any language — Qwen3-VL translates and converts it
Auto-format: Generates optimized WAN 2.2 prompt format
Multi-reference: Qwen3-VL sees all connected images via
image+image2inputsVisual style detection: 12+ artistic styles (photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy, etc.)
Smart caching: Performance optimization with Fixed Seed Mode
GGUF backend: Efficient local model inference with quantization support
Qwen3.5 support: Thinking mode disabled via
/no_thinkfor fast prompt generationCamera tag dropdown: 19 camera movements selectable directly in the node UI
🎵 MMAudio Integration
MMAudio can be added to any workflow by connecting the MMAudio nodes to the generated video output. The node analyzes the video and produces synchronized audio (music, speech, sound effects).
🎨 NSFW Support
Comprehensive content generation without restrictions
Dedicated NSFW presets for each workflow type
Natural progression, style adaptation, consistent characters
🎯 QwenVL-Mod NSFW Presets
The workflows include built-in NSFW presets for the Qwen3-VL prompt enhancer:
🍿 T2V Presets
🍿 Wan 2.2 NSFW T2V— Standard T2V prompt🍿 Wan 2.2 NSFW T2V Timeline (5s)— Timeline format for Story workflows
🎥 I2V Presets
🎥 Wan 2.2 NSFW I2V Scene (5s)— Single scene, 5 seconds📖 Wan 2.2 NSFW I2V Scene (20s)— Single scene, 20 seconds🎬 Wan 2.2 NSFW I2V Timeline (20s)— Multi-prompt timeline, 20 seconds
🔄 FL2V Presets
🔄 Wan 2.2 NSFW FL2V Scene (5s)— Transition between first and last frame
🖼️ Utility Presets
🖼️ Detailed Description— SFW detailed scene description (for non-NSFW use)
SFW presets are also available. Edit the preset dropdown in the QwenVL node to switch.
🖼️ Multi-Reference Input (image2)
The QwenVL-Mod node has two image inputs:
T2V: no images needed
I2V:
image= first frameFL2V:
image= first frame,image2= last frameSVI:
image= primary reference,image2= additional references (batch)Story:
image= first frame for I2V segments,image2= optional second reference
Qwen3-VL sees all connected images as individual images, enabling proper multi-reference analysis.
🎮 Usage Examples
Basic Text-to-Video (T2V)
Load
WAN2.2-T2V-Qwen3.5.jsonWrite your prompt in any language
Select preset
🍿 Wan 2.2 NSFW T2VGenerate video
Image-to-Video (I2V)
Load
WAN2.2-I2V-Qwen3.5.jsonUpload your first-frame image to
imageSelect preset
🎥 Wan 2.2 NSFW I2V Scene (5s)Write what happens next (in any language)
Generate animated video
First-Last-Frame (FL2V)
Load
WAN2.2-FL2V-Qwen3.5.jsonUpload first-frame to
image, last-frame toimage2Select preset
🔄 Wan 2.2 NSFW FL2V Scene (5s)Describe the transition between the two frames
Generate the interpolated video with TensorRT upscale + RIFE
Story / Timeline (I2V Story)
Load
WAN2.2-I2V-20s-Story-Qwen3.5.jsonUpload first-frame to
imageSelect preset
🎬 Wan 2.2 NSFW I2V Timeline (20s)Write prompts for each timeline segment (up to 4 prompts, 5s each)
Generate a 20-second video with automatic scene transitions
Recommended:
max_tokens = 2048,context_length = 16384+for 20s timelines
20-Second Single Scene (I2V 20s)
Load
WAN2.2-I2V-20s-Qwen3.5.jsonUpload first-frame to
imageSelect preset
📖 Wan 2.2 NSFW I2V Scene (20s)Write what happens next (in any language)
Generate a single-scene 20-second video
SVI — Subject Video Identity (20s)
Load
WAN2.2-I2V-SVI-20s-Qwen3.5.jsonUpload primary reference to
image, additional references toimage2Select preset
🎥 Wan 2.2 NSFW I2V Scene (20s)Generate a 20-second video with locked character identity
Story SVI — Timeline with Identity Lock (20s)
Load
WAN2.2-I2V-SVI-20s-Story-Qwen3.5.jsonUpload primary reference to
image, additional references toimage2Select preset
� Wan 2.2 NSFW I2V Timeline (20s)Write prompts for each timeline segment
Generate a 20-second Story video with consistent character identity
🔧 Technical Specifications
⚡ Performance
Output: 720p/1080p, 16 fps (native), up to 20 seconds (Story)
Upscale: TensorRT RealESRGAN (FL2V workflow)
Frame interpolation: RIFE v4.25 → 48 fps (FL2V workflow)
Sage Attention: FP16 accumulation, async offload
Smart caching: Reuse prompts with same inputs, Fixed Seed Mode for text-only caching
🎨 Model Support
Qwen3-VL 4B: 7 GGUF variants (2.38 GB – 4.28 GB)
Qwen3-VL 8B: 7 GGUF variants (4.8 GB – 8.71 GB)
Qwen3.5: 4B / 9B / 27B (uncensored, heretic, unsloth) — thinking mode disabled
HF Models: Josiefed, official, Heretic-Stable variants
Quantization: Q4_K_S, Q5_K_S, FP16, INT8, FP8
🌐 Multilingual Capabilities
Input languages: Any language supported
Auto-translation: Automatic translation to optimized English
Style detection: Works with multilingual prompts
Cultural adaptation: Context-aware prompt enhancement
📦 Installation
Quick Install
Download: ComfyUI-QwenVL-Mod (latest version)
Extract to
ComfyUI/custom_nodes/ComfyUI-QwenVL-ModInstall requirements:
pip install -r requirements.txtRestart ComfyUI
Load included workflows from
wan22/folder
Custom Nodes Required
ComfyUI-QwenVL-Mod — All workflows (Qwen3-VL prompt enhancer) — huchukato/ComfyUI-QwenVL-Mod
ComfyUI-RIFE-TensorRT-Auto — FL2V (frame interpolation) — huchukato/ComfyUI-RIFE-TensorRT-Auto
ComfyUI-Upscaler-TensorRT-Auto — FL2V (upscaling) — huchukato/ComfyUI-Upscaler-TensorRT-Auto
ComfyUI-VideoHelperSuite — All workflows (VHS_VideoCombine) — Kosinkadink/ComfyUI-VideoHelperSuite
ComfyUI-Easy-Use — FL2V (easy showAnything) — yolain/ComfyUI-Easy-Use
ComfyUI-PerfectVideoResolution — All workflows (resolution selector) — huchukato/ComfyUI-PerfectVideoResolution
ComfyUI-WanMoeKSampler — Story workflows (WanMoe advanced sampling) — stduhpf/ComfyUI-WanMoeKSampler
ComfyUI-PainterI2V — Story workflows (Painter I2V) — princepainter/ComfyUI-PainterI2V
ComfyUI-PainterLongVideo — Story workflows (Painter long video) — princepainter/ComfyUI-PainterLongVideo
ComfyUI-mxToolkit — Story workflows (mxSlider) — Smirnov75/ComfyUI-mxToolkit
ComfyUI-TagComplete — Wildcards (WildcardProcessor +
__pmp/prmpt/*) — huchukato/comfy-tagcompletergthree-comfy — Power Lora Loader, Fast Groups Bypasser — rgthree/rgthree-comfy
Euler-Smea-Dyn-Sampler — Alternative samplers — Koishi-Star/Euler-Smea-Dyn-Sampler
Models Required
FP8 Workflows (T2V):
models/diffusion_models/→wan22RemixT2VI2V_t2vHighV20.safetensors(~14.3 GB) orwan22RemixT2VI2V_t2vLowV20.safetensors— huchukato/garagemodels/text_encoders/→nsfw_wan_umt5-xxl_fp8_scaled.safetensors(~4.8 GB) — NSFW-API/NSFW-Wan-UMT5-XXLmodels/vae/→wan_2.1_vae.safetensors(~253 MB) — Comfy-Org
FP8 Workflows (I2V / FL2V / SVI / Story):
models/diffusion_models/→wan22RemixT2VI2V_i2vHighV30.safetensors(~14.3 GB) orwan22RemixT2VI2V_i2vLowV30.safetensors— huchukato/garageSame text encoder + VAE as T2V
TensorRT Engines (FL2V only):
models/upscale_models/→RealESRGAN_x4(TensorRT engine)models/rife/→rife425_ensemble_False_scale_1_sim(TensorRT engine)
TensorRT engines must be built for your specific GPU. See ComfyUI-RIFE-TensorRT-Auto and ComfyUI-Upscaler-TensorRT-Auto for build instructions.
Download Links
Diffusion (FP8, T2V High): WAN REMIX T2V v2.0 High — Civitai
Diffusion (FP8, T2V Low): WAN REMIX T2V v2.0 Low — Civitai
Diffusion (FP8, I2V High): WAN REMIX v2.1 FP8 High — Civitai
Diffusion (FP8, I2V Low): WAN REMIX v2.1 FP8 Low — Civitai
Text encoder (NSFW FP8): nsfw_wan_umt5-xxl_fp8_scaled.safetensors
Text encoder (standard FP8): umt5_xxl_fp8_e4m3fn_scaled.safetensors
🎬 WAN 2.2 Prompting Notes
How to Write Your Prompt
Describe the scene naturally. Be clear about the concepts below — Qwen3-VL handles the rest:
🎨 Visual style (put it first):
photorealistic,cinematic,anime,3D CG,claymation,vintage film,watercolor,fantasy👥 Subjects: number, gender, appearance, clothing, position, expression
🏃 Action / motion: what happens, speed, interaction
🎥 Camera: dolly, pan, zoom, static, handheld, crane, orbit — smooth and continuous
🌍 Environment: setting, lighting, atmosphere, time of day
🔊 Audio (optional): connect MMAudio nodes to add synchronized sound
🔄 FL2V: Describe the transition between frames, not the scene (images fix the scene) 📖 Story: Write separate prompts for each timeline segment — Qwen3-VL handles the transitions
Resolution Guidance
WAN 2.2 native resolutions:
📱 Portrait: 832×1216 · 720×1280
⬛ Square: 1024×1024
🖥️ Landscape: 1216×832 · 1280×720
⚠️ Match the aspect ratio to your input image! Forcing 16:9 on a portrait image will squash it.
Duration
Standard: 5 seconds (81 frames at 16 fps)
Story/Timeline: up to 20 seconds (4 × 5s segments)
Frame interpolation: RIFE doubles framerate to 48 fps where applicable
🎥 Camera Control Tags
All WAN 2.2 NSFW presets support camera control via the camera_tag dropdown on the QwenVL node — no need to type tags manually. Select from 19 camera movements:
[STATIC_CAMERA]/[LOCKED_OFF]— Camera completely static[SLOW_ZOOM_IN]— Slow continuous push-in[SLOW_ZOOM_OUT]— Slow continuous pull-back[FAST_ZOOM_IN]— Fast aggressive push-in[FAST_ZOOM_OUT]— Fast pull-back, reveal context[PAN_LEFT]/[PAN_RIGHT]— Smooth horizontal pan[TILT_UP]/[TILT_DOWN]— Smooth vertical tilt[DOLLY_IN]/[DOLLY_OUT]— Physical dolly movement (parallax)[TRACKING_LEFT]/[TRACKING_RIGHT]— Lateral tracking shot[CRANE_UP]/[CRANE_DOWN]— Crane/jib movement[ORBIT]— Smooth 360-degree orbit around subject[HANDHELD]— Subtle handheld sway with micro-movements[ROLL]— Slow camera roll (rotation around lens axis)
How it works: the selected tag is injected at the start of the prompt AND as a FINAL CAMERA DIRECTIVE at the end, so Qwen respects it despite recency bias. The subject stays alive and active — the tag controls only the camera.
🎲 Wildcards
Selected workflows include a WildcardProcessor node that injects randomized prompt fragments from the PMP's Prompt Engine (__pmp/prmpt/*) wildcard library.
How It Works
The WildcardProcessor node sits before the Qwen3-VL prompt enhancer
At queue time, each
__wildcard__token is replaced with a random line from the corresponding.txtfileThe expanded text is passed to Qwen3-VL, which converts it into the WAN 2.2 prompt format
Different seed = different wildcard picks — use a fixed seed for reproducible results
Customizing Wildcards
Edit existing: open the
.txtfiles underComfyUI/custom_nodes/comfy-tagcomplete/wildcards/pmp/prmpt/Add your own: create a new
.txtfile, e.g.pmp/prmpt/mytags.txt, then reference it as__pmp/prmpt/mytags__Remove a wildcard: delete the
__...__token from the WildcardProcessortextfieldDisable randomization: replace the
__wildcard__token with a fixed string
Required Custom Node
ComfyUI-TagComplete (includes the
WildcardProcessornode and the__pmp/prmpt/*wildcard set) — huchukato/comfy-tagcomplete
The wildcard files ship with the custom node as fallback. On Docker/Vast.ai deployments, wildcards are downloaded from ComfyUI-Garage at boot for the latest version.
🐳 Docker / Cloud Ready
OneClick RunPod Template
Prefer a ready-to-go environment? Use the OneClick - ComfyUI - WAN 2.2 - Qwen3VL RunPod template:
Docker image:
huchukato/comfyui-qwenvl-runpod:cu13-wan22(CUDA 13.0) orhuchukato/comfyui-qwenvl-runpod:cu128-wan22(CUDA 12.8)Base:
huchukato/comfyui-base:cu130All custom nodes pre-installed
ComfyUI Args:
--disable-auto-launch --fast fp16_accumulation --use-sage-attention --cuda-malloc --async-offloadAll 8 workflows auto-downloaded at boot
Models auto-downloaded at first boot (~62 GB including 4 WAN Remix diffusion models, NSFW text encoder, VAE; persistent)
ComfyUI v0.34.2 baked into base image
Sage Attention, FP16 accumulation, async offload
TensorRT upscaling + RIFE interpolation
PMP wildcards auto-downloaded from Garage at boot
Access: ComfyUI
:8188· JupyterLab:8888· FileBrowser:8080(useradmin/ passwordadminadmin12) · SSHssh root@pod-ip
Vast.ai Provisioning
A Vast.ai provisioning script is also available:
Script:
vastai/wan22-provisioning.shDownloads all models, workflows, wildcards, and custom nodes on first boot
Same model set as RunPod Docker
ComfyUI Args (pre-configured)
--disable-auto-launch
--fast fp16_accumulation
--use-sage-attention
--cuda-malloc
--async-offload
🚀 Why Choose ComfyUI-QwenVL-Mod + WAN 2.2?
🎬 For Content Creators
Multilingual: Write in any language, Qwen3-VL handles translation
Story/Timeline: Multi-prompt timelines for long-form content (up to 20s)
Quality: Native resolution, TensorRT upscale to higher resolution
🔥 For NSFW Content
Explicit: Uncensored generation with dedicated NSFW presets
Multiple presets: T2V, I2V (5s/20s), FL2V, Timeline — each tuned for its mode
Detailed: Rich scene descriptions with explicit action
Natural: Realistic progression, consistent characters
⚡ For Power Users
Customizable: Easy to modify presets and system prompts
Extendable: Add your own Qwen3-VL models (GGUF or HF)
Optimized: Sage Attention, FP16, async offload, smart caching
Multi-reference:
image2input for FL2V and SVI workflowsStory: WanMoeKSampler + PainterI2V for complex multi-scene generation
🌟 What Makes This Special?
Complete: 8 workflows covering T2V, I2V, FL2V, SVI, and Story
Auto-prompting: Qwen3-VL handles prompt enhancement in any language
Timeline: Multi-prompt Story workflows for up to 20-second videos
TensorRT: Built-in upscaling and frame interpolation
NSFW presets: Dedicated presets for each workflow type
Wildcards: PMP prompt engine for randomized variation
Docker-ready: OneClick RunPod template + Vast.ai provisioning
📋 Credits
WAN 2.2 — Wan-AI · Comfy-Org/Wan_2.2
ComfyUI — comfyanonymous/ComfyUI
QwenVL-Mod — huchukato/ComfyUI-QwenVL-Mod
Qwen3-VL — Qwen Team / Alibaba
WAN Remix FP8 — FX_FeiHou
NSFW Text Encoder — NSFW-API/NSFW-Wan-UMT5-XXL
TensorRT RIFE / Upscaler — huchukato
VideoHelperSuite — Kosinkadink
Easy-Use — yolain
PerfectVideoResolution — huchukato
WanMoeKSampler — stduhpf
PainterI2V / PainterLongVideo — princepainter
mxToolkit — Smirnov75
rgthree-comfy — rgthree
📄 License
Workflows are released under the same license as the underlying models and custom nodes. See each repository for details.
WAN 2.2 model weights: Wan-AI — Apache 2.0.
Built with ❤️ for the ComfyUI community
Description
FAQ
Comments (134)
Hey, thank you for your work. I noticed that llama.cpp does not seem to work with Cuda 12.9 installed. Can you confirm? The QwenVL node silently falls back to CPU and generating a prompt takes a very long time.
the wheels are for 12.8 or 13.0 but you can compile it for your hardware, for Windows with CUDA (for example) $env:CMAKE_ARGS = "-DGGML_CUDA=on" pip install "llama-cpp-python @ git+https://github.com/JamePeng/llama-cpp-python.git"
@huchukato Thank you! ❤️
QwenVL-Mod is not in the repository. You used to have a link in the description on GITHUB. Not right now. The comfyui manager does not search.
It is in the Manager, QwenVL-Mod: Enhanced Vision-Language
BTW you are right I removed the github link coz I was too happy to be in the Manager, I was supposed to leave also the link xDDD
Add “IP adapter for the face,” and you'll be counted among the saints.
Noted for the next update xD
Just a heads up - seems your RunPod container doesn't include a terminal in the Juypeter Lab setup, so awkward pulling anything outside of Civit e.g SVI loras etc. Good job though.
Thanks for let me know this I will take a look later
@huchukato No worries!
@CrackOut I just noticed that in the Connect tab there is "Enable Web Terminal" that let you use the RunPod integrated WebTerminal in the browser, BTW in the next build I added A custom webterminal on port 8081, you will find it in the Connect tab <3
@huchukato Ah yes, I've not used the webterminal before, usually used the Juypeter one opened in each file to quickly run a wget command for huggingface whilst using the civic node for the odd hosted model to play with if needed. I'll take a look at the webterminal now, I got by using a HF-Downloader haha!
The T2V/I2V WF is outstanding. I've put the standard KJ FP8 models back into the beginning to get rid of the bias of smoothmix and remix. I'd be interested in seeing a T2V/I2V SVI that can handle a full minute of video and the autoprompt to support it. Excellent work. (Running on 5090 FE, 3.13 cu130, py2.9.1)
Hi, this looks like an amazing workflow, but I am unfortunately facing the exact same issue as other commenters.
When I use the "qwen3-vl gguf" node (AILab_QwenVL GGUF) within the Autoprompt group, the node successfully triggers the automatic download of both the GGUF and mmproj files. However, even after the download is complete, I consistently get the following error:
AILab_QwenVL_GGUF_Advanced: Failed to load model from file: D:\AI\StabilityMatrix-win-x64\Data\Packages\ComfyUI\models\LLM\GGUF\mradermacher\Qwen3-VL-8B-Instruct-c_abliterated-v3-GGUF\Qwen3-VL-8B-Instruct-c_abliterated-v3.Q4_K_M.gguf
I have tried restarting ComfyUI multiple times, but the error persists. I've spent nearly 12 hours troubleshooting this with Gemini and ChatGPT, but nothing has worked so far. It seems like a persistent bug on Windows environments.
I am using Windows 11 and Stability Matrix. I suspect it might be related to the long file path or how the node handles GGUF loading on Windows. I would truly appreciate it if you could look into this bug.
you tried another model? The 4B or another quantization of the 8B?
@huchukato Yes.
In the model_name (in the qwen3-vl gguf node), I tried:
qwen3-vl-8b-instruct-c_abliterated-v3.q4_k_m.gguf
qwen3-vl-8b-instruct-c_abliterated-v3.q6_k.gguf
qwen3-vl-4b-instruct-c_abliterated-v2.q4_k_m.gguf
However, I encountered the same error with all of them.
And following Gemini’s advice, I installed CUDA, the NVIDIA CUDA Toolkit, and Microsoft Visual Studio. I also showed the ComfyUI CMD (console) window to Gemini and ChatGPT, but there was no problem and the result was the same.
Also the manager inside ComfyUI did not detect any missing files. It seems that all the required files are installed.
@huchukato Also, I have a question. In the "qwen3-vl autoprompt" group, there are four prompts (prompt1–4). I noticed that I can’t directly edit the text written inside them. Is that how it’s supposed to be?
(I’m referring to the text that says “two woman with vibrant green eyes~”.)
(I already have the “comfyui-easy-use” node.)
The QwenVL GGUF nodes need llama-cpp-python with vision support. A normal pip install llama-cpp-python often doesn’t include handlers.. Install the vision-capable llama-cpp-python
python -m pip install --upgrade --no-cache-dir --force-reinstall ` "llama-cpp-python @ git+https://github.com/JamePeng/llama-cpp-python.git"
@drernestbrown474 After following your instructions and reinstalling llama-cpp-python with vision support, qwen3-vl seems to be working now. Thank you for that!
@twopoints225221 the prompts are generated by Qwen following your prompt and the image, that's why I called it "OneClick", you can also try just with the image and no prompt and Qwen will generate a story following the initial reference image
@huchukato I didn’t know that—what an incredible feature!
I’m using it well, but I have a question.
When writing a prompt, how can I make the 'qwen3-vl gguf' node structure the prompt by time segments?
For example, how would I give a command like, “From 0 to 5 seconds, walk and from 5 to 10 seconds, run”?
@twopoints225221 if you use the story wf just try Prompt 1: (something) prompt 2 etc, you you use the full i2v autoprompt, you can write a different prompt every 5 seconds by your own and it generates 1 prompt every 5 seconds of video
@huchukato Thank you!
I see you are the author of the qwen-vl-mod for comfyui. As there is no possibility to make an issue on your github: can you please add ltx2 prompt presets? because now it will make prompts for 5 seconfs, but ltx2 can make 20 secs....
Ok, sorry but I am daft. I don't understand how the Qwen3-VL Autoprompt is supposed to work in the workflow. What is the format for what I am supposed to add to the text field. I disable it because it generates things in the prompt that I did not ask for.
t. A simillar work flow called Smooth uses florence2run. Init's case to try to make video from a picture and the ai tries to figure out what'd be going on. In it's case you can disable those nodes manually and describe action yourself. The also did a all in one wokflow kind of thing where you can pick what you want to do. One I grabbed didn't pick up on what ever the author has set up for dependencies, and my reaction is: I have better things to do then fuck around with this.
I am currently using OneClick-I2v-Svi-Story, and it is working very well — thank you for the great workflow!
Ran into issues with the Qwen-VL GGUF Advanced node where it wasn't cleaning up vram correctly and would fail after 2 runs.
https://github.com/1038lab/ComfyUI-QwenVL/issues/104
Ended up swapping the GGUF py file from your repo for the ones here and it hasn't been an issue since. However I do lose all the additions you made to your fork.
It's a great workflow and runs perfectly on my 4070. Takes about 18 minutes for a full 20 second clip at 640x960, upscaled 30fps, FP8.
having the following error when i try to run this
SyntaxError: Unexpected non-whitespace character after JSON at position 4 (line 1 column 5)
which WF?
I encountered the same error when trying the OneClick-I2V-Story workflow. The latest version of the ComfyUI desktop app is currently 0.14.1 (it doesn't automatically update, even though 0.14.2 is out). I encountered the same issue with that version, and the error disappeared when I unpacked all subgraphs. Try unpacking all subgraphs.
According to ChatGPT, if you're not using the latest ComfyUI version, something is causing the issue with the subgraphs.
@fopof4264449 what do you mean by "unpacked" all subgraphs?
Guys I'm using 0.14.2
@huchukato Me too. Didn't fiddle too much, got the same error message.
Very nice enjoyable workflow! The autoprompt works great. I wish there was any way to avoid the quality drop in later half of longer videos though, it feel inevitable
I have to work on the SVI workflow for that but I got issues with slow motion
@huchukato Yeah I experience same, longer videos loses its point when one have to speed it up , hence again creating a relatively short video😅 I hope you figure out something clever🙌🏻
so this can use taek's fp8 nsfw low/high v2? would you recommend high motion or the normal version for realistic scenes?
I would really appreciate a little help.
I have installed everything, followed every step, downloaded the neccessary stuff. I tried to use the mmost simple I2V workflow, without promp help, but my MoE KSampler just refuses to work. It doesn't give me an error, just highlighted in red, and the proccess wont go above like 6%.
Any ideas?
I have the same issue, it lights up red almost instantly when generation starts, the prompt isn't even generated and still it's red, and in console I can see at least "Failed to validate prompt for output..." error.
I spent a moment more looking into why this doesn't work in my case;
* WanMoeKSamplerAdvanced ...
- Return type mismatch between linked nodes: scheduler, received_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent']) mismatch input_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent', 'beta57'])
Output will be ignored
Looks like at least in my case there's something changed somewhere, and the enum doesn't anymore match - there's beta57 in another of the enums, which comes from RES4LYF I think. But if change the sampler to be manually configured within the subgraph, at least this error/issue is gone.
the scheduler issue is due to RES4LYF
how would i setup Sageattention to speed things along? or do you have worflows with it already?
@huchukato
also, where the negative prompt box?
Hello, I tried to "go with the normal version Full-I2V-LongVideo-v1.7". Yet I am still prompted by ComfyUI to install the QWEN-stuff. Is the workflow in the file correct?
With the last update this WF contains both normal and GGUF Long Video WFs coz I made too many variants and I was not able to manage them anymore, so I grouped them. To use the normal Qwen3-VL node you have to load WAN2.2-I2V-Full-AutoPrompt-MMAudio-v1-9.json from here https://civitai.com/models/2320999?modelVersionId=2613591
The scheduler selection causes currently issues, which stops the video subgraphs from functioning, specifically the Wan MoE KSampler, as it expects different enum (missing at least one scheduler name, in my case beta57. There is a mismatch in the enum with latest ComfyUI node versions of the used custom nodes as of today (23.2)
I don't have beta57 in latest ComfyUI 0.14.2, maybe you installed it with RES4LYF?
I tryed to install RES4LYF and it brokes everything coz the KSampler does not recognize the scheduler, if you want to use that node you have to change the KSampler.
@huchukato Yes, it's not a problem for me, I can get around it easy enough but just thought to point it out that there's possible incompatibilities.
Same here, but I'm on 0.15 which probably is troublesome.
I wrote about RES4LYF in the readme, didn't tested 0.15 yet
@i_m_l_O how do you get around it easily? do you replace the ksampler? do I need 2 ksamplers now cause the regular clownsharksampler only got 1 model connection.
@Smurfypie I think you need 2 clownshark, I didn't tested it till now
ok cock fucker. Prompt outputs failed validation: AILab_QwenVL_GGUF_Advanced: - Value not in list: preset_prompt: '🍿 Wan 2.2 NSFW I2V' not in ['🖼️ Tags', '🖼️ Simple Description', '🖼️ Detailed Description', '🖼️ Ultra Detailed Description', '🎬 Cinematic Description', '🖼️ Detailed Analysis', '📹 Video Summary', '📖 Short Story', '🪄 Prompt Refine & Expand'] WanMoeKSamplerAdvanced: - Return type mismatch between linked nodes: scheduler, received_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent']) mismatch input_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent', 'beta57']) . what piece of shit.
@huchukato huchukato
OP
" I wrote about RES4LYF in the readme, didn't tested 0.15 yet" wtf are you even trying to say?
QwenVLGGUFBase.run() missing 1 required positional argument: 'unload_after_run' I get this error and cant seem to fix it.
I am confused about the upscalers. If I want to use tensorrt upscaler do I turn on both Enable Upscaler and Enable Upscaler TENSORRT, or just Enable Upscaler TENSORRT?
Just the Tensorrt one and link the "setupscale" setnode the the one you are using
I'm getting this error when the workflow moves onto the 10s subgraph.
AILab_QwenVL_Advanced
t:1 must be larger than temporal_factor:2
This is on the I2V-Full-AutoPrompt-MMAudio workflow. Any ideas?
I'm also stuck at "Loading checkpoint shards: 100%" when using the Qwen3-VL-8B-Instruct-Ablitered model. Waited over 20 minutes. I'm able to generate a 5s video with Qwen3-VL-8B-Instruct and Qwen3-VL-8B-Instruct-FP8 models however.
Is it supposed to build a rife engine every single time I run the workflow?
The shitstain didn't even bother to test some stuff it's just a broken cunt ass add.
Probably. the shit stain has no idea wtf he's doing
I am getting an error about OOM when using the I2V prompt workflow, specifically when using Qwen ( any model, doesn't matter if it fits GPU ):
AILab_QwenVL_Advanced
Allocation on device 0 would exceed allowed memory. (out of memory) Currently allocated : 0 bytes Requested : 1.16 GiB Device limit : 31.37 GiB Free (according to CUDA): 14.46 GiB PyTorch limit (set by user-supplied memory fraction) : 17179869184.00 GiB This error means you ran out of memory on your GPU. TIPS: If the workflow worked before you might have accidentally set the batch_size to a large number.
What the hell is going on? Has someone else had this issue?
which WF? the Story one or the Full one?
@huchukato Specifically WAN2.2-I2V-AutoPrompt.json, it's not story or full, I just want the basic Image to Video, but it keeps throwing that weird memory error.
@pipindirovskigorjo318 i check and let you know ASAP
Updated both normal and GGUF Single Video WFs, download them again and let me know https://civitai.com/models/2320999?modelVersionId=2624175
@huchukato Same error unfortunately :
[QwenVL] Loading Qwen3-VL-4B-Instruct-Abliterated (8-bit (Balanced), attn=sdpa)
Loading weights: 0%| | 0/713 [00:00<?, ?it/s]
!!! Exception during processing !!! Allocation on device 0 would exceed allowed memory. (out of memory)
Currently allocated : 9.12 MiB
Requested : 741.88 MiB
Device limit : 31.37 GiB
Free (according to CUDA): 22.41 GiB
PyTorch limit (set by user-supplied memory fraction)
: 17179869184.00 GiB
Traceback (most recent call last):
File "/ComfyUI/execution.py", line 524, in execute
output_data, output_ui, has_subgraph, has_pending_tasks = await get_output_data(prompt_id, unique_id, obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/ComfyUI/execution.py", line 333, in get_output_data
return_values = await asyncmap_node_over_list(prompt_id, unique_id, obj, input_data_all, obj.FUNCTION, allow_interrupt=True, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/ComfyUI/execution.py", line 307, in asyncmap_node_over_list
await process_inputs(input_dict, i)
File "/ComfyUI/execution.py", line 295, in process_inputs
result = f(**inputs)
^^^^^^^^^^^
File "/ComfyUI/custom_nodes/ComfyUI-QwenVL-Mod/AILab_QwenVL.py", line 706, in process
return self.run(model_name, quantization, preset_prompt, custom_prompt, image, video, frame_count, max_tokens, temperature, top_p, num_beams, repetition_penalty, seed, keep_model_loaded, attention_mode, use_torch_compile, device, keep_last_prompt)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/ComfyUI/custom_nodes/ComfyUI-QwenVL-Mod/AILab_QwenVL.py", line 585, in run
self.load_model(
File "/ComfyUI/custom_nodes/ComfyUI-QwenVL-Mod/AILab_QwenVL.py", line 468, in load_model
self.model = AutoModelForVision2Seq.from_pretrained(model_path, **load_kwargs).eval()
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/comfyui-env/lib/python3.12/site-packages/transformers/models/auto/auto_factory.py", line 374, in from_pretrained
return model_class.from_pretrained(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/comfyui-env/lib/python3.12/site-packages/transformers/modeling_utils.py", line 4137, in from_pretrained
loading_info, disk_offload_index = cls._load_pretrained_model(model, state_dict, checkpoint_files, load_config)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/comfyui-env/lib/python3.12/site-packages/transformers/modeling_utils.py", line 4256, in loadpretrained_model
loading_info, disk_offload_index = convert_and_load_state_dict_in_model(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/comfyui-env/lib/python3.12/site-packages/transformers/core_model_loading.py", line 1212, in convert_and_load_state_dict_in_model
realized_value = mapping.convert(
^^^^^^^^^^^^^^^^
File "/opt/comfyui-env/lib/python3.12/site-packages/transformers/core_model_loading.py", line 678, in convert
collected_tensors = self.materialize_tensors()
^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/comfyui-env/lib/python3.12/site-packages/transformers/core_model_loading.py", line 654, in materialize_tensors
tensors = [future.result() for future in tensors if future.result() is not None]
^^^^^^^^^^^^^^^
File "/usr/lib/python3.12/concurrent/futures/_base.py", line 449, in result
return self.__get_result()
^^^^^^^^^^^^^^^^^^^
File "/usr/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result
raise self._exception
File "/usr/lib/python3.12/concurrent/futures/thread.py", line 58, in run
result = self.fn(*self.args, **self.kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/comfyui-env/lib/python3.12/site-packages/transformers/core_model_loading.py", line 800, in _job
return materializecopy(tensor, device, dtype)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/comfyui-env/lib/python3.12/site-packages/transformers/core_model_loading.py", line 789, in materializecopy
tensor = tensor.to(device=device, dtype=dtype)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
torch.OutOfMemoryError: Allocation on device 0 would exceed allowed memory. (out of memory)
Currently allocated : 9.12 MiB
Requested : 741.88 MiB
Device limit : 31.37 GiB
Free (according to CUDA): 22.41 GiB
PyTorch limit (set by user-supplied memory fraction)
: 17179869184.00 GiB
Memory summary: |===========================================================================|
| PyTorch CUDA memory summary, device ID 0 |
|---------------------------------------------------------------------------|
| CUDA OOMs: 0 | cudaMalloc retries: 0 |
|===========================================================================|
| Metric | Cur Usage | Peak Usage | Tot Alloc | Tot Freed |
|---------------------------------------------------------------------------|
| Allocated memory | 83589 KiB | 8418 MiB | 0 B | 0 B |
| from large pool | 0 KiB | 0 MiB | 0 B | 0 B |
| from small pool | 0 KiB | 0 MiB | 0 B | 0 B |
|---------------------------------------------------------------------------|
| Active memory | 83589 KiB | 8418 MiB | 0 B | 0 B |
| from large pool | 0 KiB | 0 MiB | 0 B | 0 B |
| from small pool | 0 KiB | 0 MiB | 0 B | 0 B |
|---------------------------------------------------------------------------|
| Requested memory | 0 B | 0 B | 0 B | 0 B |
| from large pool | 0 B | 0 B | 0 B | 0 B |
| from small pool | 0 B | 0 B | 0 B | 0 B |
|---------------------------------------------------------------------------|
| GPU reserved memory | 131072 KiB | 8480 MiB | 0 B | 0 B |
| from large pool | 0 KiB | 0 MiB | 0 B | 0 B |
| from small pool | 0 KiB | 0 MiB | 0 B | 0 B |
|---------------------------------------------------------------------------|
| Non-releasable memory | 0 B | 0 B | 0 B | 0 B |
| from large pool | 0 B | 0 B | 0 B | 0 B |
| from small pool | 0 B | 0 B | 0 B | 0 B |
|---------------------------------------------------------------------------|
| Allocations | 0 | 0 | 0 | 0 |
| from large pool | 0 | 0 | 0 | 0 |
| from small pool | 0 | 0 | 0 | 0 |
|---------------------------------------------------------------------------|
| Active allocs | 0 | 0 | 0 | 0 |
| from large pool | 0 | 0 | 0 | 0 |
| from small pool | 0 | 0 | 0 | 0 |
|---------------------------------------------------------------------------|
| GPU reserved segments | 0 | 0 | 0 | 0 |
| from large pool | 0 | 0 | 0 | 0 |
| from small pool | 0 | 0 | 0 | 0 |
|---------------------------------------------------------------------------|
| Non-releasable allocs | 0 | 0 | 0 | 0 |
| from large pool | 0 | 0 | 0 | 0 |
| from small pool | 0 | 0 | 0 | 0 |
|---------------------------------------------------------------------------|
| Oversize allocations | 0 | 0 | 0 | 0 |
|---------------------------------------------------------------------------|
| Oversize GPU segments | 0 | 0 | 0 | 0 |
|===========================================================================|
@pipindirovskigorjo318 after a loooong night of debugging I finally fixed it xD Download the WF again and update my Qwen Node (just do an Update All in ComfyUI Manager) and it will work
i'm not crashing with an OOM message, but my 5090 runs full, and then it just stops processing, the 4B or 8B model doesn't matter.
@malur You are running the last WFs version and the last Qwen node version? coz I fixed that
@huchukato i downloaded them literally earlier today, after some more testing, i guess only the 8B-MAX VL Version doesn't like to run consistently. fresh launch and it runs, next run it just just stops or takes ages, idk, i stopped it after like 10 minutes last time. can i somehow convert it to gguf or nvfe4 to save on some vram? i mean, 8B is huge anyway.
@malur mmm I disable the quantization in the normal node coz it caused a lot of problems so in that node all the models runs only in FP16, if you wanna use the quantized models you have to go for the GGUF node and install llama cpp
Looking very much to trying these workflows. Do you have a Url for the StoryTeller repository? The one I googled is no longer available.
Thanks
I noticed that some parameters in Oneclick and fulllongvideo seem to be misaligned. For example, it shows max token 0, temperature 512, top p 0.6, but normally it should be max token 512, temperature 0.6, and so on. The ComfyUI version is 0.16.3. I tried switching versions, but the issue persists. The OneVideo workflow parameters are normal and functioning correctly.
Hello, i dont understand how im supposed to install qwen LV model, can someone help me please ?
Just select a model in the node and it will auto download it
One day ill find a work flow that works for me, every single one always has a point where something will just refuse the install, this time its tensorrt
Why does any model loaded to qwinVL node if it is Q4 it gives me uncensored response but Q8 never give me uncensored response?
Found a bug in qwenvl-mod: NSFW content is missing from the generated prompts. I found on GitHub that the AILab_System_Prompts.json file was modified on Mar 5, 2026, and the NSFW-related descriptions were removed from the prompt.
Just got home and added the NSFW prompt back in—tested it and everything works fine now.
I checked the preset now and there is
WHEN there are NSFW images or text, provide an NSFW description consistent with the requested artistic styleas always, I just changed the presets names
Ok I see the problem, the extended NSFW check was only in 1 preset. Fixed the others right now
Now we just need a "OneClick-T2V I2V SVI -Story
I planned to make it but I got slow motion issues with SVI and I was trying to fix that issues before
@huchukato make sure to hardcode more repos that are half baked in while your at it. oh and more cunt-um nodes. I don't think people are trolled enough make sure it never garabage collects to while you vibe code.
@bugsbelightyear730 Man, if you don't like what I did, just don't use it. Easy peasy.
when i try to use any kind of wan2.2 lora
it seems to turn into a blurry mess, had to adjust a few settings to keep moving people from becoming blurry. however couldn t get loras working. are there specific lora type to be used for this?
Having the same issue
Depends on which Wan 2.2 Model are you running
your step size is too low. You need to adjust your stepsize to higher like 20, or use a lightning LORA
You should add a referral code for Vast.AI and RunPod, happy to have used it!
WE'LL TEST TO SEE AND THEN I'LL MAKE A RETURN
There are issues and conflicts with this setup and the nodes it's looking for, at least on Linux. I can't fully explain them because I had to poke at it a bit but first I thought there was an issue with the ~/models/llm directory because the NSFW Qwen LLM wasn't showing up and something is looking for /LLM and something else is looking in /llm which are separate directories in Linux but not in Windows. Then I realized that if /ComfyUI-QwenVL exists then the node picks that instead of the intended /ComfyUI-QwenVL-Mod. So I had to remove /ComfyUI-QwenVL and then the node showed the expected gguf listings. I would assume this issue would have been recognized in the runpod setup since that is Linux, I'm running it locally on a Linux box.
UNFORTUNATELY, NOTHING IS WORKING PROPERLY. WHAT A SHAME.
If you think you came here to find workflows to get running WAN right now, you came to the wrong place. At least at this moment in time. All of these are currently impossibly broken.
shits taking 7min to forge a 5sec video on a 5080 on 832*480 res. no tutorial no nothing
you speak about motion frames in imtovid Description: Number of trailing frames from previous_video used as motion reference
Default: 5 where isi it? i don't understand that ? can you help me?
I keep getting image token error. I believe all the nodes are up to date. Anyone have any advice on the issue?
[load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842]
Using I2V AutoPrompt Story and it's throwing TypeError: VRAMCleanup.empty_cache() got an unexpected keyword argument 'input'. Swap out to LayerUtility VRAMCleanup and same error.
because the stupid guy on the node 5s / 10s / 15s / 20s . Is sending wan moe decode latent *purplee dot to > vram cleanup input.
When he should be sending latent to > wan vae decode > samples.
from there send wan vae decode output to image
delete vram cleanup.
Because i dont know where in the hell i should connect that
also disconnect the node in wan moe decode > sample name > just leave it without the green cable. Otherwise will thow error of mismatch.
and now it works. Repeat this in all others nodes, meaning 10s 15s 20s.
@messinajonathan2459 Thank you for the solution. I delete the vramcleanup node and it works fine. Still a wonderful workflow.
For anyone experiencing Qwen prompt issues on the "Story" workflows. Qwen3-VL-4B will not be able to understand the Timeline (20s) instructions. Had to switch to a quantized 8B model for it to understand and generate the prompts correctly.
is totally broken given errors of mismatch
[QwenVL] GPU memory after load: 16.3GB / 31.8GB
got prompt
Failed to validate prompt for output 1327:
* VAEDecode 1252:1658:
- Return type mismatch between linked nodes: samples, received_type(*) mismatch input_type(LATENT)
Output will be ignored
Failed to validate prompt for output 1252:1657:
* (prompt):
- Return type mismatch between linked nodes: input, received_type(LATENT) mismatch input_type(*)
* VRAMCleanup 1252:1657:
- Return type mismatch between linked nodes: input, received_type(LATENT) mismatch input_type(*)
Output will be ignored
Failed to validate prompt for output 1314:
* AutoRifeTensorrt 1596:
- Return type mismatch between linked nodes: frames, received_type(*) mismatch input_type(IMAGE)
Output will be ignored
Failed to validate prompt for output 1608:
* (prompt):
- Return type mismatch between linked nodes: input, received_type(IMAGE) mismatch input_type(*)
* VRAMCleanup 1608:
- Return type mismatch between linked nodes: input, received_type(IMAGE) mismatch input_type(*)
Output will be ignored
Failed to validate prompt for output 1596:
* (prompt):
- Return type mismatch between linked nodes: frames, received_type(*) mismatch input_type(IMAGE)
Output will be ignored
Failed to validate prompt for output 1547:
Output will be ignored
got prompt
Failed to validate prompt for output 1327:
* VAEDecode 1252:1658:
- Return type mismatch between linked nodes: samples, received_type(*) mismatch input_type(LATENT)
Output will be ignored
Failed to validate prompt for output 1252:1657:
* (prompt):
- Return type mismatch between linked nodes: input, received_type(LATENT) mismatch input_type(*)
* VRAMCleanup 1252:1657:
- Return type mismatch between linked nodes: input, received_type(LATENT) mismatch input_type(*)
Output will be ignored
Failed to validate prompt for output 1314:
* AutoRifeTensorrt 1596:
- Return type mismatch between linked nodes: frames, received_type(*) mismatch input_type(IMAGE)
Output will be ignored
Failed to validate prompt for output 1608:
* (prompt):
- Return type mismatch between linked nodes: input, received_type(IMAGE) mismatch input_type(*)
* VRAMCleanup 1608:
- Return type mismatch between linked nodes: input, received_type(IMAGE) mismatch input_type(*)
Output will be ignored
Failed to validate prompt for output 1596:
* (prompt):
- Return type mismatch between linked nodes: frames, received_type(*) mismatch input_type(IMAGE)
Output will be ignored
Failed to validate prompt for output 1547:
Output will be ignored
got prompt
Failed to validate prompt for output 1327:
* VAEDecode 1252:1658:
- Return type mismatch between linked nodes: samples, received_type(*) mismatch input_type(LATENT)
Output will be ignored
Failed to validate prompt for output 1547:
Output will be ignored
Failed to validate prompt for output 1252:1657:
* (prompt):
- Return type mismatch between linked nodes: input, received_type(LATENT) mismatch input_type(*)
* VRAMCleanup 1252:1657:
- Return type mismatch between linked nodes: input, received_type(LATENT) mismatch input_type(*)
the fucking stupid guy sent wan moe decoder latent to vram clean up
when you should be sending latent to > van decode samples.
vram cleanup is fucked up or something is missing
is nsfw blocked in this flow? Can't seem to get any nsfw content
auto prompt works fine, but no video is generated. i don't know what i am doing wrong, no errors i receive
Same. Just keep spitting out empty video files and two png's.
i discover that i have some issues with sage, i bypass the node and woks fine now.
@dantufis664 Good to know. Thanks!
The workflow works great, but I am having an issue with the Qwen3VL nodes - they seem to be running on CPU even though I have cuda selected as the device to use. Takes a very long time to load each prompt. Any known fix?
I watched the tutorial video and followed it as it was, but most of the models didn't install, and I couldn't. Just wasted time and runpod credits.
Using your WAN2.2-I2V-AutoPrompt with 60FPS enabled, I get this error:
File "/ComfyUI/execution.py", line 308, in _async_map_node_over_list await process_inputs(input_dict, i) File "/ComfyUI/execution.py", line 296, in process_inputs result = f(**inputs) ^^^^^^^^^^^ File "/ComfyUI/custom_nodes/ComfyUI-RIFE-TensorRT-Auto/__init__.py", line 420, in load_rife_tensorrt_model engine.build( File "/ComfyUI/custom_nodes/ComfyUI-RIFE-TensorRT-Auto/trt_utilities.py", line 338, in build raise RuntimeError("TensorRT not available - please install TensorRT first")
I installed requirements_cu12.txt to fix it, but I feel like I shouldn't have really done that.
Did I miss a config setting, picked a wrong pod? Don't want to reinstall again the next time
The battle with TensorRT lasted half a day :) But I emerged victorious. I'll describe this using the RTX5060ti drv 595.79 torch as an example: 2.10.0+cu130 Python 3.13.11 ComfyUI 0.19.4
The problem is that it's looking for a dll in a folder that doesn't exist - ComfyUI_windows_portable\python_embeded\Lib\site-packages\tensorrt\tensorrt.libs
But simply creating it doesn't help.
1. Download the one you need from https://developer.nvidia.com/tensorrt/download/10x
2. Unzip the TensorRT-10.16.1.11.Windows.amd64.cuda-13.2.zip archive to G:\TensorRT). You should now have a TensorRT folder containing python, lib, bin, etc. The python folder contains the required whl (do not use the dispatch and lean versions).
3. Install from ComfyUI_windows_portable via CMD - .\python_embeded\python.exe -m pip install G:\TensorRT\python\tensorrt-10.16.1.11-cp313-none-win_amd64.whl - substitute the required .whl here.
4. Create a tensorrt.libs folder in the G:\ComfyUI_windows_portable\python_embeded\Lib\site-packages\tensorrt folder.
5. Copy all files from the bin folder of the TensorRT-10.16.1.11.Windows.amd64.cuda-13.2.zip archive to the folder. G:\ComfyUI_windows_portable\python_embeded\Lib\site-packages\tensorrt\tensorrt.libs
Open the Environment Variables window:
Press Win + R, type sysdm.cpl, and press Enter. Go to the "Advanced" → "Environment Variables..." tab. Find and edit the Path variable: Under "System Variables," double-click Path, click New
And paste the path:
G:\ComfyUI_windows_portable\python_embeded\Lib\site-packages\tensorrt\tensorrt.libs
Save the changes. Click "OK" in all open windows. Afterwards, be sure to restart your computer for the changes to take effect. Done! TensorRT is installed. Now restart ComfyUI. The ComfyUI-RIFE-TensorRT-Auto plugin should load without errors, and the AutoLoadRifeTensorrtModel and AutoRifeTensorrt nodes will become available.
Nice work!
But after the update, I can't select model or preset in the Qwen Nord...
i can absolutely not install llama-cpp
what is the comfyui virtual environment?
Really great, the autoprompt works very well. I tried replacing the PainterI2V node with the wanFirstLastFrameTo video node and accurately describing the action, there was a great transition between the two photos except in the last frames with the usual problems of aligning the character with the final reference photo, but only for the eyes. With this method, however, Qwen-VL obviously only analyzes the starting photo. Would it be possible to have Qwen-VL analyze both the starting and final images? It would be great if it automatically created a prompt for the transition between two photos.
cant get the tensorrt nodes to work, what are alternatives?
From user swdvb:
"How to install TensorRT
I use ComfyUI on Windows11.
ComfyUI version: 0.3.68
Python version: 3.12.10
pytorch version: 2.8.0+cu128
I read all your comments and the methods each of you used, but none of you specified exactly what needs to be done to make everything work.
I used the hints (ai-capybara), but that’s not the whole process.
What needs to be done:
Download tensorrt-10.13.2.6 from the NVidia website: (https://developer.nvidia.com/tensorrt/download/10x)
Choose your Python version (mine is Python version: 3.12.10). To install, open the folder with (python_embeded), right-click and choose (open in terminal) from the menu, then enter these commands one by one:
( python.exe -s -m pip install wheel-stub )
( python.exe -m pip install --upgrade pip setuptools wheel )
( python.exe -m pip install nvidia-cudnn-cu12==9.2.0.82 onnx polygraphy --no-cache-dir )
Once these libraries are installed, proceed to install tensorrt.
If you downloaded the file, for example (tensorrt-10.13.2.6-cp312-none-win_amd64.whl),
and it’s located in another folder or on another drive, you don’t need to copy it to the python folder—just enter this command ( python.exe -m pip install ) and then drag and drop the file into the terminal window and press Enter.
If the file is located in the python folder, then run this command:
( python.exe -m pip install tensorrt-10.13.2.6-cp312-none-win_amd64.whl )
Just copy the name of the file you want to install and replace (tensorrt-10.13.2.6-cp312-none-win_amd64.whl) with your version.
After all this, you need to copy all dll files from the folder (TensorRT-10.13.2.6\lib) to the python_embeded directory.
That’s the whole process. After that, everything should start up. If these nodes don’t work (ComfyUI-Upscaler-Tensorrt) and (ComfyUI_TensorRT), you need to install (numpy-1.26.4.dist-info). How to install it:
Go to (https://pypi.org/project/numpy/1.26.4/#files) and download the (whl) file for your Python version and install it the same way as tensorrt—just enter this command ( python.exe -m pip install ) and then drag and drop the file into the terminal window and press Enter."
Thank you!! Greatly appreciated on the model links inside comfyui
hey, thanks for your work, i noticed that you have some workflows with mmaudio but anyone with SVI and mmaudio, is this compatible?
