CivArchive

    โœจOne-click Pod available on:โœจ

    ๐ŸŸฃ Deploy on RunPod with CUDA 13.0

    ๐ŸŸฃ Deploy on RunPod with CUDA 12.8

    ๐ŸŸก Deploy on VastAI

    ๐Ÿณ RunPod users: Just click the template link, choose a GPU, and everything installs automatically โ€” ComfyUI, all nodes, all workflows, and WAN 2.2 models (~30GB) download in the background on first boot. No manual setup needed. ComfyUI starts immediately while models download.


    โ˜•๏ธ buymeacoffee


    IMPORTANT:

    If you install RES4LYF node it will broke the MoEKSampler, to use it you have to use the KSampler included in that node.


    ComfyUI-QwenVL-Mod โ€” Enhanced Vision-Language with WAN 2.2 Version 2.9.0 (2026/09/07) โ€” ๐ŸŽฌ WAN 2.2 NSFW Video + WAN Remix T2V/I2V Models + Story/Timeline Workflows (up to 20s) + SVI Camera + FL2V First-Last-Frame + 8 Workflows + Wildcards Included


    โฌ†๏ธ 2026/09/07 UPDATE โฌ†๏ธ

    ๐Ÿ“ฆ What's Included โ€” 8 Workflows

    All workflows are pre-wired with Qwen3-VL auto-prompting, WAN Remix diffusion models, and TensorRT upscale + RIFE interpolation where applicable.

    • WAN2.2-T2V-Qwen3.5.json โ€” T2V ยท Text-to-video, 5 seconds

    • WAN2.2-I2V-Qwen3.5.json โ€” I2V ยท Image-to-video, 5 seconds

    • WAN2.2-FL2V-Qwen3.5.json โ€” FL2V ยท First-Last-Frame to video, TensorRT upscale + RIFE

    • WAN2.2-I2V-20s-Qwen3.5.json โ€” I2V 20s ยท Single-scene image-to-video, 20 seconds

    • WAN2.2-I2V-20s-Story-Qwen3.5.json โ€” Story I2V ยท Multi-prompt timeline, 20 seconds (4 ร— 5s)

    • WAN2.2-I2V-SVI-20s-Qwen3.5.json โ€” SVI 20s ยท Subject Video Identity, 20 seconds

    • WAN2.2-I2V-SVI-20s-Story-Qwen3.5.json โ€” Story SVI ยท Timeline with SVI identity lock, 20 seconds

    • WAN2.2-T2V-I2V-Story-Qwen3.5.json โ€” Story T2V+I2V ยท Timeline mixing T2V and I2V, 20 seconds


    ๐Ÿ”„ WAN Remix T2V/I2V Models โ€” 4 Variants

    All WAN 2.2 workflows now use the WAN Remix T2V/I2V diffusion models. Download from the original Civitai pages:

    ๐Ÿงน Removed: WAN Enhanced NSFW SVI Camera

    • Removed wan22EnhancedNSFWSVICamera_nsfwV2FP8H/L models โ€” superseded by WAN Remix

    • Docker and provisioning cleaned up

    ๐ŸŽฒ PMP Wildcards โ€” Downloaded at Boot

    • Wildcards (__pmp/prmpt/*) are now downloaded from ComfyUI-Garage at boot time

    • No Docker rebuild needed to update wildcards โ€” just push to Garage and restart the pod

    • comfy-tagcomplete ships with wildcard fallback for local installs


    โฌ†๏ธ 2026/08/04 UPDATE โฌ†๏ธ

    โœจ ComfyUI QwenVL-Mod Node Update โœจ

    v2.4 โ€” Local Model Discovery + Qwen3.5 + SageAttention

    We haven't forgotten about this node! Here's what's new since v2.2:

    • ๐ŸŽฌ LTX 2.3 Presets (v2.3): New specialized presets for LTX 2.3 I2V and T2V with official prompting guides. Multilingual support for all presets, simplified single-paragraph format (max 200 words), full NSFW support.

    • ๐Ÿ” Local Model Discovery (v2.4): Drop your GGUF/HF files into models/LLM/ and they show up in the dropdown automatically โ€” no more JSON editing. Auto-pairs mmproj files for vision GGUF models.

    • ๐Ÿง  Qwen3.5 support (v2.4): Architecture detection from file metadata (GGUF header / HF config.json), automatic thinking-mode disabling, forced top_k=20.

    • โšก SageAttention restored (v2.4): Architecture-aware kernels (Blackwell FP8, Hopper FP8, Ada FP8, Ampere FP16) with graceful SDPA fallback.

    ๐Ÿ”ง Also updated our companion nodes:

    • ComfyUI-Upscaler-TensorRT-Auto โ€” TensorRT upscaling with auto-detection, CUDA 12/13 wheels pre-baked

    • ComfyUI-RIFE-TensorRT-Auto โ€” TensorRT frame interpolation, CUDA 12/13 wheels pre-baked

    • comfy-tagcomplete โ€” Tag completion with wildcard support for WAN 2.2 workflows

    • ComfyUI-HuggingFace โ€” Model download integration for local discovery

    All WAN 2.2 and LTX 2.3 workflows (T2V, I2V, SVI, MMAudio, GGUF variants) are tested and working with the updated nodes. Grab the latest version and let us know how it goes!


    โš ๏ธ Requirements โ€” Read First!

    GPU & VRAM

    • ๐ŸŸข Recommended โ€” RTX 5090 (32 GB) / RTX PRO 6000 (48 GB) / RTX 4090 (24 GB) โ†’ FP8 Remix models

    • ๐ŸŸก Mid-range โ€” RTX 3090 (24 GB) / RTX 4080 (16 GB) โ†’ FP8 with offload

    • ๐ŸŸ  Lower VRAM โ€” 12โ€“16 GB โ†’ FP8 with aggressive offload

    Model Quantization Options

    • FP8 (recommended) โ€” ~14.3 GB per diffusion model + ~4.8 GB text encoder = ~19 GB active set โ†’ huchukato/garage

    • FP16 (full) โ€” ~42 GB per diffusion model + ~12 GB text encoder = ~54 GB total โ†’ Comfy-Org/Wan_2.2

    Software

    • ComfyUI: v0.31.0+

    • Python: 3.10+

    • CUDA: 12.8+ (13.0 recommended)

    • Storage: allow at least 80 GB for the complete provisioned package

    Text Encoder

    • FP8 (recommended, NSFW): nsfw_wan_umt5-xxl_fp8_scaled.safetensors (~4.8 GB) โ€” NSFW-API/NSFW-Wan-UMT5-XXL

    • FP8 (standard): umt5_xxl_fp8_e4m3fn_scaled.safetensors โ€” Comfy-Org


    ๐ŸŒŸ What is ComfyUI-QwenVL-Mod?

    A powerful enhanced vision-language node for ComfyUI that combines Qwen3-VL models with WAN 2.2 video generation workflows. Features multilingual support, visual style detection, NSFW capabilities, Story/Timeline multi-prompt generation, and MMAudio integration.

    Think: "Your all-in-one solution for intelligent prompt enhancement and video generation with WAN 2.2!"


    ๐ŸŽฌ Key Features

    ๐Ÿš€ WAN 2.2 Video Generation

    • T2V (Text-to-Video): Generate video from text prompts

    • I2V (Image-to-Video): Animate a first-frame image

    • FL2V (First-Last-Frame): Generate the transition between two keyframes โ€” Qwen3-VL sees both frames

    • SVI (Subject Video Identity): Lock character identity across generations using reference images

    • Story (Timeline): Multi-prompt timeline generation โ€” up to 4 prompts for 20-second videos with automatic scene transitions

    ๐Ÿง  Qwen3-VL Auto-Prompting

    • Multilingual: Write your prompt in any language โ€” Qwen3-VL translates and converts it

    • Auto-format: Generates optimized WAN 2.2 prompt format

    • Multi-reference: Qwen3-VL sees all connected images via image + image2 inputs

    • Visual style detection: 12+ artistic styles (photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy, etc.)

    • Smart caching: Performance optimization with Fixed Seed Mode

    • GGUF backend: Efficient local model inference with quantization support

    • Qwen3.5 support: Thinking mode disabled via /no_think for fast prompt generation

    • Camera tag dropdown: 19 camera movements selectable directly in the node UI

    ๐ŸŽต MMAudio Integration

    MMAudio can be added to any workflow by connecting the MMAudio nodes to the generated video output. The node analyzes the video and produces synchronized audio (music, speech, sound effects).

    ๐ŸŽจ NSFW Support

    • Comprehensive content generation without restrictions

    • Dedicated NSFW presets for each workflow type

    • Natural progression, style adaptation, consistent characters


    ๐ŸŽฏ QwenVL-Mod NSFW Presets

    The workflows include built-in NSFW presets for the Qwen3-VL prompt enhancer:

    ๐Ÿฟ T2V Presets

    • ๐Ÿฟ Wan 2.2 NSFW T2V โ€” Standard T2V prompt

    • ๐Ÿฟ Wan 2.2 NSFW T2V Timeline (5s) โ€” Timeline format for Story workflows

    ๐ŸŽฅ I2V Presets

    • ๐ŸŽฅ Wan 2.2 NSFW I2V Scene (5s) โ€” Single scene, 5 seconds

    • ๐Ÿ“– Wan 2.2 NSFW I2V Scene (20s) โ€” Single scene, 20 seconds

    • ๐ŸŽฌ Wan 2.2 NSFW I2V Timeline (20s) โ€” Multi-prompt timeline, 20 seconds

    ๐Ÿ”„ FL2V Presets

    • ๐Ÿ”„ Wan 2.2 NSFW FL2V Scene (5s) โ€” Transition between first and last frame

    ๐Ÿ–ผ๏ธ Utility Presets

    • ๐Ÿ–ผ๏ธ Detailed Description โ€” SFW detailed scene description (for non-NSFW use)

    SFW presets are also available. Edit the preset dropdown in the QwenVL node to switch.


    ๐Ÿ–ผ๏ธ Multi-Reference Input (image2)

    The QwenVL-Mod node has two image inputs:

    • T2V: no images needed

    • I2V: image = first frame

    • FL2V: image = first frame, image2 = last frame

    • SVI: image = primary reference, image2 = additional references (batch)

    • Story: image = first frame for I2V segments, image2 = optional second reference

    Qwen3-VL sees all connected images as individual images, enabling proper multi-reference analysis.


    ๐ŸŽฎ Usage Examples

    Basic Text-to-Video (T2V)

    1. Load WAN2.2-T2V-Qwen3.5.json

    2. Write your prompt in any language

    3. Select preset ๐Ÿฟ Wan 2.2 NSFW T2V

    4. Generate video

    Image-to-Video (I2V)

    1. Load WAN2.2-I2V-Qwen3.5.json

    2. Upload your first-frame image to image

    3. Select preset ๐ŸŽฅ Wan 2.2 NSFW I2V Scene (5s)

    4. Write what happens next (in any language)

    5. Generate animated video

    First-Last-Frame (FL2V)

    1. Load WAN2.2-FL2V-Qwen3.5.json

    2. Upload first-frame to image, last-frame to image2

    3. Select preset ๐Ÿ”„ Wan 2.2 NSFW FL2V Scene (5s)

    4. Describe the transition between the two frames

    5. Generate the interpolated video with TensorRT upscale + RIFE

    Story / Timeline (I2V Story)

    1. Load WAN2.2-I2V-20s-Story-Qwen3.5.json

    2. Upload first-frame to image

    3. Select preset ๐ŸŽฌ Wan 2.2 NSFW I2V Timeline (20s)

    4. Write prompts for each timeline segment (up to 4 prompts, 5s each)

    5. Generate a 20-second video with automatic scene transitions

    6. Recommended: max_tokens = 2048, context_length = 16384+ for 20s timelines

    20-Second Single Scene (I2V 20s)

    1. Load WAN2.2-I2V-20s-Qwen3.5.json

    2. Upload first-frame to image

    3. Select preset ๐Ÿ“– Wan 2.2 NSFW I2V Scene (20s)

    4. Write what happens next (in any language)

    5. Generate a single-scene 20-second video

    SVI โ€” Subject Video Identity (20s)

    1. Load WAN2.2-I2V-SVI-20s-Qwen3.5.json

    2. Upload primary reference to image, additional references to image2

    3. Select preset ๐ŸŽฅ Wan 2.2 NSFW I2V Scene (20s)

    4. Generate a 20-second video with locked character identity

    Story SVI โ€” Timeline with Identity Lock (20s)

    1. Load WAN2.2-I2V-SVI-20s-Story-Qwen3.5.json

    2. Upload primary reference to image, additional references to image2

    3. Select preset ๏ฟฝ Wan 2.2 NSFW I2V Timeline (20s)

    4. Write prompts for each timeline segment

    5. Generate a 20-second Story video with consistent character identity


    ๐Ÿ”ง Technical Specifications

    โšก Performance

    • Output: 720p/1080p, 16 fps (native), up to 20 seconds (Story)

    • Upscale: TensorRT RealESRGAN (FL2V workflow)

    • Frame interpolation: RIFE v4.25 โ†’ 48 fps (FL2V workflow)

    • Sage Attention: FP16 accumulation, async offload

    • Smart caching: Reuse prompts with same inputs, Fixed Seed Mode for text-only caching

    ๐ŸŽจ Model Support

    • Qwen3-VL 4B: 7 GGUF variants (2.38 GB โ€“ 4.28 GB)

    • Qwen3-VL 8B: 7 GGUF variants (4.8 GB โ€“ 8.71 GB)

    • Qwen3.5: 4B / 9B / 27B (uncensored, heretic, unsloth) โ€” thinking mode disabled

    • HF Models: Josiefed, official, Heretic-Stable variants

    • Quantization: Q4_K_S, Q5_K_S, FP16, INT8, FP8

    ๐ŸŒ Multilingual Capabilities

    • Input languages: Any language supported

    • Auto-translation: Automatic translation to optimized English

    • Style detection: Works with multilingual prompts

    • Cultural adaptation: Context-aware prompt enhancement


    ๐Ÿ“ฆ Installation

    Quick Install

    1. Download: ComfyUI-QwenVL-Mod (latest version)

    2. Extract to ComfyUI/custom_nodes/ComfyUI-QwenVL-Mod

    3. Install requirements: pip install -r requirements.txt

    4. Restart ComfyUI

    5. Load included workflows from wan22/ folder

    Custom Nodes Required

    Models Required

    FP8 Workflows (T2V):

    • models/diffusion_models/ โ†’ wan22RemixT2VI2V_t2vHighV20.safetensors (~14.3 GB) or wan22RemixT2VI2V_t2vLowV20.safetensors โ€” huchukato/garage

    • models/text_encoders/ โ†’ nsfw_wan_umt5-xxl_fp8_scaled.safetensors (~4.8 GB) โ€” NSFW-API/NSFW-Wan-UMT5-XXL

    • models/vae/ โ†’ wan_2.1_vae.safetensors (~253 MB) โ€” Comfy-Org

    FP8 Workflows (I2V / FL2V / SVI / Story):

    • models/diffusion_models/ โ†’ wan22RemixT2VI2V_i2vHighV30.safetensors (~14.3 GB) or wan22RemixT2VI2V_i2vLowV30.safetensors โ€” huchukato/garage

    • Same text encoder + VAE as T2V

    TensorRT Engines (FL2V only):

    • models/upscale_models/ โ†’ RealESRGAN_x4 (TensorRT engine)

    • models/rife/ โ†’ rife425_ensemble_False_scale_1_sim (TensorRT engine)

    TensorRT engines must be built for your specific GPU. See ComfyUI-RIFE-TensorRT-Auto and ComfyUI-Upscaler-TensorRT-Auto for build instructions.


    ๐ŸŽฌ WAN 2.2 Prompting Notes

    How to Write Your Prompt

    Describe the scene naturally. Be clear about the concepts below โ€” Qwen3-VL handles the rest:

    • ๐ŸŽจ Visual style (put it first): photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy

    • ๐Ÿ‘ฅ Subjects: number, gender, appearance, clothing, position, expression

    • ๐Ÿƒ Action / motion: what happens, speed, interaction

    • ๐ŸŽฅ Camera: dolly, pan, zoom, static, handheld, crane, orbit โ€” smooth and continuous

    • ๐ŸŒ Environment: setting, lighting, atmosphere, time of day

    • ๐Ÿ”Š Audio (optional): connect MMAudio nodes to add synchronized sound

    ๐Ÿ”„ FL2V: Describe the transition between frames, not the scene (images fix the scene) ๐Ÿ“– Story: Write separate prompts for each timeline segment โ€” Qwen3-VL handles the transitions

    Resolution Guidance

    WAN 2.2 native resolutions:

    • ๐Ÿ“ฑ Portrait: 832ร—1216 ยท 720ร—1280

    • โฌ› Square: 1024ร—1024

    • ๐Ÿ–ฅ๏ธ Landscape: 1216ร—832 ยท 1280ร—720

    โš ๏ธ Match the aspect ratio to your input image! Forcing 16:9 on a portrait image will squash it.

    Duration

    • Standard: 5 seconds (81 frames at 16 fps)

    • Story/Timeline: up to 20 seconds (4 ร— 5s segments)

    • Frame interpolation: RIFE doubles framerate to 48 fps where applicable

    ๐ŸŽฅ Camera Control Tags

    All WAN 2.2 NSFW presets support camera control via the camera_tag dropdown on the QwenVL node โ€” no need to type tags manually. Select from 19 camera movements:

    • [STATIC_CAMERA] / [LOCKED_OFF] โ€” Camera completely static

    • [SLOW_ZOOM_IN] โ€” Slow continuous push-in

    • [SLOW_ZOOM_OUT] โ€” Slow continuous pull-back

    • [FAST_ZOOM_IN] โ€” Fast aggressive push-in

    • [FAST_ZOOM_OUT] โ€” Fast pull-back, reveal context

    • [PAN_LEFT] / [PAN_RIGHT] โ€” Smooth horizontal pan

    • [TILT_UP] / [TILT_DOWN] โ€” Smooth vertical tilt

    • [DOLLY_IN] / [DOLLY_OUT] โ€” Physical dolly movement (parallax)

    • [TRACKING_LEFT] / [TRACKING_RIGHT] โ€” Lateral tracking shot

    • [CRANE_UP] / [CRANE_DOWN] โ€” Crane/jib movement

    • [ORBIT] โ€” Smooth 360-degree orbit around subject

    • [HANDHELD] โ€” Subtle handheld sway with micro-movements

    • [ROLL] โ€” Slow camera roll (rotation around lens axis)

    How it works: the selected tag is injected at the start of the prompt AND as a FINAL CAMERA DIRECTIVE at the end, so Qwen respects it despite recency bias. The subject stays alive and active โ€” the tag controls only the camera.


    ๐ŸŽฒ Wildcards

    Selected workflows include a WildcardProcessor node that injects randomized prompt fragments from the PMP's Prompt Engine (__pmp/prmpt/*) wildcard library.

    How It Works

    1. The WildcardProcessor node sits before the Qwen3-VL prompt enhancer

    2. At queue time, each __wildcard__ token is replaced with a random line from the corresponding .txt file

    3. The expanded text is passed to Qwen3-VL, which converts it into the WAN 2.2 prompt format

    4. Different seed = different wildcard picks โ€” use a fixed seed for reproducible results

    Customizing Wildcards

    • Edit existing: open the .txt files under ComfyUI/custom_nodes/comfy-tagcomplete/wildcards/pmp/prmpt/

    • Add your own: create a new .txt file, e.g. pmp/prmpt/mytags.txt, then reference it as __pmp/prmpt/mytags__

    • Remove a wildcard: delete the __...__ token from the WildcardProcessor text field

    • Disable randomization: replace the __wildcard__ token with a fixed string

    Required Custom Node

    The wildcard files ship with the custom node as fallback. On Docker/Vast.ai deployments, wildcards are downloaded from ComfyUI-Garage at boot for the latest version.


    ๐Ÿณ Docker / Cloud Ready

    OneClick RunPod Template

    Prefer a ready-to-go environment? Use the OneClick - ComfyUI - WAN 2.2 - Qwen3VL RunPod template:

    • Docker image: huchukato/comfyui-qwenvl-runpod:cu13-wan22 (CUDA 13.0) or huchukato/comfyui-qwenvl-runpod:cu128-wan22 (CUDA 12.8)

    • Base: huchukato/comfyui-base:cu130

    • All custom nodes pre-installed

    • ComfyUI Args: --disable-auto-launch --fast fp16_accumulation --use-sage-attention --cuda-malloc --async-offload

    • All 8 workflows auto-downloaded at boot

    • Models auto-downloaded at first boot (~62 GB including 4 WAN Remix diffusion models, NSFW text encoder, VAE; persistent)

    • ComfyUI v0.34.2 baked into base image

    • Sage Attention, FP16 accumulation, async offload

    • TensorRT upscaling + RIFE interpolation

    • PMP wildcards auto-downloaded from Garage at boot

    Access: ComfyUI :8188 ยท JupyterLab :8888 ยท FileBrowser :8080 (user admin / password adminadmin12) ยท SSH ssh root@pod-ip

    Vast.ai Provisioning

    A Vast.ai provisioning script is also available:

    • Script: vastai/wan22-provisioning.sh

    • Downloads all models, workflows, wildcards, and custom nodes on first boot

    • Same model set as RunPod Docker

    ComfyUI Args (pre-configured)

    --disable-auto-launch
    --fast fp16_accumulation
    --use-sage-attention
    --cuda-malloc
    --async-offload
    

    ๐Ÿš€ Why Choose ComfyUI-QwenVL-Mod + WAN 2.2?

    ๐ŸŽฌ For Content Creators

    • Multilingual: Write in any language, Qwen3-VL handles translation

    • Story/Timeline: Multi-prompt timelines for long-form content (up to 20s)

    • Quality: Native resolution, TensorRT upscale to higher resolution

    ๐Ÿ”ฅ For NSFW Content

    • Explicit: Uncensored generation with dedicated NSFW presets

    • Multiple presets: T2V, I2V (5s/20s), FL2V, Timeline โ€” each tuned for its mode

    • Detailed: Rich scene descriptions with explicit action

    • Natural: Realistic progression, consistent characters

    โšก For Power Users

    • Customizable: Easy to modify presets and system prompts

    • Extendable: Add your own Qwen3-VL models (GGUF or HF)

    • Optimized: Sage Attention, FP16, async offload, smart caching

    • Multi-reference: image2 input for FL2V and SVI workflows

    • Story: WanMoeKSampler + PainterI2V for complex multi-scene generation


    ๐ŸŒŸ What Makes This Special?

    • Complete: 8 workflows covering T2V, I2V, FL2V, SVI, and Story

    • Auto-prompting: Qwen3-VL handles prompt enhancement in any language

    • Timeline: Multi-prompt Story workflows for up to 20-second videos

    • TensorRT: Built-in upscaling and frame interpolation

    • NSFW presets: Dedicated presets for each workflow type

    • Wildcards: PMP prompt engine for randomized variation

    • Docker-ready: OneClick RunPod template + Vast.ai provisioning


    ๐Ÿ“‹ Credits


    ๐Ÿ“„ License

    Workflows are released under the same license as the underlying models and custom nodes. See each repository for details.

    WAN 2.2 model weights: Wan-AI โ€” Apache 2.0.


    Built with โค๏ธ for the ComfyUI community

    Description

    Removed the Tensorrt Upscaler and replaced with 2xLexicaRRDBNet

    FAQ

    Comments (41)

    xuf129750620Jan 26, 2026
    CivitAI

    AILab_QwenVL_GGUF_Advanced

    [QwenVL] llama_cpp is not available. Install the GGUF vision dependency first. See docs/GGUF_MANUAL_INSTALL.md

    huchukato
    Author
    Jan 26, 2026ยท 1 reaction

    I wrote the guide to install it here and also in the WF...

    meritrash6350Jan 26, 2026

    @huchukatoย Yeah,the guide isn't geared toward everyone. You might have come up with something cool here, but the explanations on how to get llama going are extremely lacking.

    huchukato
    Author
    Jan 27, 2026

    @meritrash6350ย I wrote all the steps ypu have to do to install llama, the only thing that lack is how to activate the virtual enviroment on Windows coz I don't use Windows, I'm on Mac and Linux

    meritrash6350Jan 27, 2026

    @huchukatoย So, you get my point then?

    huchukato
    Author
    Jan 28, 2026

    @meritrash6350ย I got it but you don't have to install the GGUF version of the WF in every cases, go with the normal one where you don't have to install llama python, I cannot tell you how to get in a venv in Windows not having Windows on my PC :\

    evantopsmithJan 28, 2026ยท 1 reaction

    @meritrash6350 In your ComfyUI root installation folder activate enviornment using either script:

    Command Prompt: \venv\Scripts\activate.bat

    PowerShell: \venv\Scripts\Activate.ps1

    meritrash6350Jan 28, 2026

    @evantopsmithย Sorry, wasn't ignoring you, I got too busy. I'll take a look at it, and I appreciate the extra effort. I personally find the effort of dealing with venv to be the thing I hate most about Python. That and the fact it insists on caching on the C: even though that's not where I told it to install.

    huchukato
    Author
    Jan 28, 2026

    @evantopsmithย Thanks a lot

    lanceshockerJan 26, 2026
    CivitAI

    I am so confused on what you mean to install. The guide isn't clear, whad you you even mean by start a ComfyUI virtual environment???

    huchukato
    Author
    Jan 27, 2026

    When you install ComfyUI the installer create a virtual enviroment to run it, usually thers is a "venv" folder inside ComfyUI. This allows Comfy to run on a specific Python version and to install all the dependecies that it needs not for all you system but just for Comfy, to avoid comflicts in you system

    qwe246Jan 27, 2026
    CivitAI

    Hello, in my QWENVL node, the generated prompt are completely unrelated to the prompt I entered. The generated prompt seem to only describe the content of the image itself. Why is this happening? Full i2v longvideo gguf

    Adam_LolFeb 2, 2026

    I am having the same issue

    evantopsmithJan 28, 2026
    CivitAI

    I am about to lose my fucking mind with this Qwen3 VL GGUF auto prompt bullshit. git cloned QwenVL-Mod (no QwenVL node from manager) start ComfyUI install dependencies, quit. download llama_cpp_python-0.3.23+cu128.basic-cp312-win_amd64.whl (Python 3.12.10, CUDA 2.8.0cu128) cmd in python_embedded (equivalent to being in active virtual environment using comfyui-easy-install) pip install --upgrade --force-reinstall llama_cpp_python-0.3.23+cu128.basic-cp312-win_amd64.whl (whl file is in same folder cmd was started in) restart comfyui and try to use Huihui-Qwen3-VL-4B-Instruct-abliterated-Q8_0.gguf.....

    AILab_QwenVL_GGUF_Advanced

    [QwenVL] Missing Qwen VL chat handler in llama_cpp. Install the correct fork/wheel. See docs/GGUF_MANUAL_INSTALL.md

    huchukato
    Author
    Jan 28, 2026

    :\

    huchukato
    Author
    Jan 28, 2026

    @gregariousbuttons257ย This is the original node but you have to manual install the abliterated Qwen models, the only difference with mine is that I added the uncensored models

    evantopsmithJan 28, 2026

    @huchukatoย but I followed all of that to the tee. which part of what I described am I doing wrong? do I have the wrong whl for my python cuda versions?

    yamyproJan 28, 2026
    CivitAI

    Was so hyped for this workflow but I have been trying to get it to work for the last 3 hours. I have yet to make even one generation! something with the Ksampers? Which I don't understand because they work normally in other workflows but for some reason this one just refuses to work for me!

    this is what i keep getting:
    Failed to validate prompt for output 1431:

    * KSamplerAdvanced 1252:1269:

    - Return type mismatch between linked nodes: scheduler, received_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent']) mismatch input_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent', 'beta57'])

    * KSamplerAdvanced 1252:1270:

    - Return type mismatch between linked nodes: scheduler, received_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent']) mismatch input_type(['simple', 'sgm_uniform', 'karras', 'exponential', 'ddim_uniform', 'beta', 'normal', 'linear_quadratic', 'kl_optimal', 'bong_tangent', 'beta57'])

    Output will be ignored

    Failed to validate prompt for output 1252:1259:

    Output will be ignored

    Failed to validate prompt for output 1327:

    Output will be ignored

    [QwenVL] Flash-Attn auto mode: dependency not ready, using SDPA

    Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592

    Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592

    Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592

    Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592

    Interrupting prompt 1cc5223a-200f-42bd-908f-13af58061592

    Processing interrupted

    Prompt executed in 524.73 seconds

    huchukato
    Author
    Jan 28, 2026

    Seems there is a problem with the setnode with the scheduler, which version are you using?

    jarigoni949Jan 28, 2026

    @huchukato, i have the same problem, what do you mean by Version? Comfy?ย 

    huchukato
    Author
    Jan 28, 2026

    @jarigoni949ย Comfy and also the WF version, coz for me all versions works on Comfy 0.10.0, Python 3.12.12 and CU13.0 with Pytorch 2.9.1

    jarigoni949Jan 29, 2026

    @huchukatoย Hey :)!
    I'am on: Comfy 0.11.0 / Python 3.12.10 / pytorch version: 2.9.1+cu130 /

    huchukato
    Author
    Jan 29, 2026

    @jarigoni949ย I'm testing the WF on 0.11.0 now and works :\ You can try to do a thing: disconnect the scheduler and sampler from the subgraph and manually set them inside it, maybe it's a problem of the Selectors node

    SolHelJan 29, 2026ยท 1 reaction

    I agree. If it's not about sampler issue, it's about QWEN abridged. The model doesn't download automatically, and when you do manually, the node won't detect it and download other models.

    yamyproJan 31, 2026

    im using comfy11.1, Python 3.12.11, 2.10.0+cu130 and CUDA13.0

    huchukato
    Author
    Jan 31, 2026

    @yamyproย The WF is tested on Python 3.12.12, update your py version maybe

    yamyproFeb 1, 2026

    @huchukatoย i know this is a dumb question but how do i upgrade it?

    jetpisces619Jan 30, 2026
    CivitAI

    I still have the error: The checkpoint you are trying to load has model type qwen3_vl but Transformers does not recognize this architecture. This could be because of an issue with the checkpoint, or because your version of Transformers is out of date.

    huchukato
    Author
    Jan 30, 2026

    You updated my Qwen node? Now it supports both transformers>5.0 and <5.0, I fixed the deprecated syntax

    redlucario1735Jan 31, 2026
    CivitAI

    Ive got it working with no major issues, the only problem that i have is that the generated prompts are very censored, it doesnt describe any nsfw behaviors and at most it implies vaguely that some "intimacy" might be going on. Any idea on how to make it as detailed and nsfw as possible?

    huchukato
    Author
    Jan 31, 2026

    If you are using my modified node and one of the Qwen Abliterated models it should work, for me it works

    dukefan6842872Jan 31, 2026
    CivitAI

    [QwenVL] Failed to apply SageAttention patch: module 'transformers.models.qwen2.modeling_qwen2' has no attribute 'F'

    [QwenVL] SageAttention patch failed, continuing with SDPA

    Is this a SageAttention // Transformers version issue?

    huchukato
    Author
    Jan 31, 2026

    you have sageattention==2.2.0? If it fail to load it it uses SDPA by default

    dukefan6842872Feb 1, 2026

    Name: sageattention

    Version: 2.2.0+cu130torch2.9.0andhigher.post4

    I do yeah, and SDPA still only takes a moment to run QwenVL, but it's still bothering me.

    dukefan6842872Feb 1, 2026ยท 1 reaction

    I've got things tuned so that this workflow is working really well for me. My main issue now is that UpscalerTensorRT at 2x upscale on a 10 second clip is using an insane amount of vram and spilling into shared memory causing it to take upwards of 30min to run, when the rest of the generation only takes ~10 minutes. (RTX 4090) Trying to figure out if it's possible to optimize UpscalerTensorRT's memory usage or if some models loaded that don't need to be when the workflow gets there.
    Edit: I moved the post-processing (UpscalerTensorRT and Interpolation) into a separate workflow and it finishes in around a minute on my video output, so there's something that I don't have the understanding to fix regarding memory management to run your entire workflow efficiently.

    dukefan6842872Feb 1, 2026

    Oh, and thanks for your responses and thanks very much for sharing the workflow!!!

    huchukato
    Author
    Feb 1, 2026

    @dukefan6842872ย no problem, btw now I'm getting the same error LOL, I will take a look tomorrow. Regarding Tensorrt, try the normal Upscale and check if it takes less time. PS. I just update again the node, now have more aderence in NSFW prompting, in minutes I also release the SVI version of the WF

    huchukato
    Author
    Feb 1, 2026

    @dukefan6842872ย I think I fixed the problem, update the node and let me know thank u

    ComfyWorkflows
    Wan Video 2.2 I2V-A14B

    Details

    Downloads
    119
    Platform
    CivitAI
    Platform Status
    Deleted
    Created
    1/26/2026
    Updated
    7/31/2026
    Deleted
    1/31/2026

    Files

    WAN22NSFWI2VT2VWorkflowsAutoPrompt_fullI2VLVGGUF.zip