Qwen 3 VL 4B quantized as w4a8.
Only tested as text encoder / clip for Krea2.
Needs minimum ComfyUI v.0.31.0
Description
FAQ
Comments (11)
[Generate Text] int8 and int4 take 16 seconds, while this int8 asym 78 seconds, does it have to be so slow? upd: a ComfyUI bug
it's not int8 actually, its w4a8 , so quantized 4 bit running at 8, also are you con comfyui v0.31.0?
@TiwazM I have already updated comfy kitchen and the back end, and I already had the newest Torch for CUDA 13
@who_is_civet I don't know w4a8 support is very new and like a day old in ComfyIi. I am using it as a text encoder / clip with no issues.
@who_is_civet are you trying to run it as an LLM ?
@TiwazM Generate Text only, so the ComfyUI devs should fix it, there is nothing to do yet. Works as intended to encode text
as of now didn't find any issue or any quality loss or bad outputs. works really well and the lower file size is a extreme plus. thank you
Not noticing any quality differences. Without extreme testing, it's only SLIGHTLY slower in my 8gb VRAM / 32gb RAM setup. May just keep for sheer size savings. :D
Benched aganist a FP8_Scaled version
@xFennec777 thanks for the feedback, I did it mostly for fun to see if it works ;)
Same as I did with one of my zimage turbo which is about SDXL size in total Unet and TE ;)
@TiwazM Oh, totally! That's what I am doing - I spent a lot of time making GGUF tooling for Flux Klein 2. I personally enjoy feedback on my work, happy to not give so much if I become too noisy on your work.
