Merge of my favourite Wan image enhancing loras, makes it much faster to use than when adding the lora separately.
Samplers
In my testing/research, I have found the best Sampler is Res_3s with Bong_Tangent Scheduler both from the RES4LYF custom nodes pack: https://github.com/ClownsharkBatwing/RES4LYF
I am using this great free txt2img workflow: https://pastebin.com/GPYQjUrx from AItrepreneur check out his Patreon https://www.patreon.com/c/aitrepreneur/posts
V2 is WAN 2.2. Merge that can do incredibly low steps down to 2 steps.
4 steps look better and 8 steps is great quality.
If you want a bit more realism, you can use the LightX lora with a negative weight, but you might have to then increase steps.
Can make decent images/videos in as low as 6 steps as it has speed loras mixed in.
Description
Initial Release
FAQ
Comments (32)
im probably stupid but i did a fresh comfyui install just for wan, then loaded the workflow, installed all the nodes and loaded up the work flow and hit run and then Ksampler select, Basic Scheduler, and Ksampler all just light up red then nothing happens. my 2nd time trying to just get wan to work and its another fail lmao
edit: tried the aitrepreneur workflow too and get "Error
Failed to get input node 0 for group node child 433:1 with slot 1" ... i love how comfy just like, never works. really hope to try wan one day lmao
Sounds like either something isn't connected properly or the samplers/schedulers used in the workflow are installed on your system - which is definitely possible as you need to separately get the RES4LYF pack.
If you don't install RES4LYF, Comfy (for the time being) won't know what res_3s and bong_tangent are...
FluxFest i got the kijai workflow at https://civitai.com/models/1818841/wan-22-workflow-t2v-i2v-t2i-kijai-wrapper to work but only after i disabled patch sage attention, model patch torch settings, and torch compile model wan lmao. its not previewing in the sampler as it generates but it eventually makes the image. im sure it takes a lot longer this way but at least im in a workable state for now lol.
cutetodeath78409597 If you don't have sage attention and triton installed and you're on Windows, I'd check out this post on reddit to save yourself a lot of headaches:
https://old.reddit.com/r/comfyui/comments/1l94ynk/so_anyways_i_crafted_a_ridiculously_easy_way_to/
Having some issues with this... As it reads like a checkpoint, I tried loading it as one and Comfy helpfully informed me that it doesn't have a clip or vae loaded... so I loaded it instead as a diffusion model and replaced the usual clip/vae nodes.
Even at low resolution settings and low frames, I'm getting crazy VRAM usage (I'm on an rtx5070 12gb vram) that I don't usually see on these res/frame settings.
I also tried it with res/2s and bong_tangent as recommended, 6 steps... but the resulting output is pixellated, blurred and a total mess despite taking about 10 minutes to generate.... any ideas?
hmm are you using the V3 FP8 version?
Maybe 6 steps at Res_2s is not enough with out some extra lighting lora.
I use Res_3s but it will take even longer.
I will play about next week and see if I can find some better settings for lower Vram GPU's.
I managed to make this image: https://civitai.com/images/93034555
with 6 steps of Res_2s but I had to put the Lightning lora: https://huggingface.co/Kijai/WanVideo_comfy/tree/main/Wan22-Lightning
all the way to 1.0 strengh to get a good image else it was very blurry @6 steps with res_2s
J1B Yeah it's the V3 FP8. Which lightning Lora is the checkpoint using?
this is awesome, thank you for your work as always!
is i wanted to train a lora, would you recommend using this as the base checkpoint or better to use base WAN and then load the lora on top of this checkpoint?
I don't think it would matter too much as it is quite close to the base model. I haven't tried it either way yet.
this is both high noise and low noise in one? which 2 loras did you add to it? instagirl and lenovo? or others?
Is this both? High noise and low noise?
so good effect, add gguf plz
Hello, what is the purpose of the ‘auto variation’ text node? I see that it adds all of its text to all prompts and does not vary. Have I misunderstood something? Thank you.
The {} syntax randomly picks a different input for each image, but you think it is not working? I have noticed difference in variation while using it, but there could be issues with the node string joining.
@J1B No, I used a "show text" and everything was copied. Maybe I made a mistake.
@Light_x02 Yeah I did some testing today and it doesn't seem to work like the person on Reddit said it did, maybe I am doing something wrong??? I will try and fix it.
Hi, this model looks impressive, I have few questions:
1. May it work with 8gb VRAM for img creation?
2. Is it intended to create only img or vid too?
3. Do you think to import it on tensor art too? It may bypass my specs limit...
Thanks.
I have created a simple but effective workflow for this model: https://civitai.com/models/1949103?modelVersionId=2205961
I will test this out, it looks good, my posted workflow is a bit complex and intimidating for ComfyUI beginners.
@J1B Thanks :) and thank you for your hard work on the WAN model. I hope you will continue working on it - it is amazing and has almost completely replaced Flux models for me (I would just wish there were more loras for WAN). Prompt following, fine detail coherence and anatomy just works better.
Hello, is the V3 indeed a Wan 2.2 model? Thank you :)
Would it be possible to create an I2V version?
This can do I2V well as well, it is just I mainly used it for images. I even had it on the generator a few weeks ago as a video model, but no one really used it.
@J1B What would an I2V workflow look like with this model?
@J1B Do you mean T2V? Your WAN SeeMe v3 is actually the best T2V model that I've ever used! Thank you for it. My only wish is that it would do I2V (tried WanImageToVideo) or if I added extra state dict VACE, it would become VACE compatible. The reason why I wish for I2V compatibility is so that I can take the T2V and then extend it with I2V. In your workflows, you have Jib I2I and T2I examples, but all your I2V WFs seem to be FusionX and not Jib.
@NiceTurtle Ahh yes I see my mistake, yes I meant Text to Video.
I could look at doing something with the separate Image to video model.
But I am quite focused on Qwen-Image at the moment, as it is a bit more flexible at Text to image than Wan and that is mainly what I use AI for right now.
make a I2V you handsome devil
Pure Gold!
I would like to re-request the bid to have a retrained version from the I2V WAN model. This is still the best T2V/T2I low noise model, but when making clips longer than 5 seconds, I have to extend with something like Remix and it loses the Jib quality. Maybe have a pre-release so that it can raise buzz donations.
