Update: WAN 2.2 Smooth Workflow v6.0 - 07/18/2026
"WAN 2.2 Smooth Workflow v6.0" update: The new version now has a First2LastFrame Workflow. Make sure to fully update your Comfyui and the extensions.
At the time of making the Showcase: Stability Matrix v2.16.1 and ComfyUI v.0.28.0.
Changed the Interpolation extension: Now using the ComfyUI-FrameInterpolation
Changed the default Resolution and added a recommended resolutions graph.
Changed the default Negative Prompt
Update a few some of the nodes
The videos on Showcase used the SmoothMix T2V v4.0
All Extensions for WAN 2.2 Smooth Workflow v6.0:
Update: WAN 2.2 Smooth Workflow v5.0 - 04/19/2026
"WAN 2.2 Smooth Workflow v5.0" update: The new version now has a First2LastFrame Workflow. Make sure to fully update your Comfyui and the extensions.
At the time of making the Showcase: Stability Matrix v2.15.7 and ComfyUI v.019.3.
Lightx2v Loras used:
HIGH with weight 1.0 -> high_noise_lora_rank64_lightx2v_4step_1022.safetensors
LOW with weight 1.5 - > lightx2v_T2V_14B_cfg_step_distill_v2_lora_rank128_bf16.safetensors
Still uses the same extensions as v4.0 - No need to install more. Make sure ComfyUI, its dependencies and the Extensions are fully updated.
The Showcase is focused on the usage of First2LastFrame since the other workflows are pretty much the same. Added some examples of Loops you can make. I might take a few tries but its possible to make something simple - keep Loops at 5 seconds at most - Wan 2.2 doesn't work very well above that threshold.
Updated Notes: Added mode hyperlinks for some of the Resources and Models Extensions used in the Workflow.
The videos on Showcase used the SmoothMix I2V v.2.
All Extensions for WAN 2.2 Smooth Workflow v5.0:
Links for MMAudio usage:
MMAudio NSFW Model (fine-tuned off the base model)
Put the MMAudio models on ComfyUI/models/mmaudio. Create the folder if it does not exist.
Update: WAN 2.2 Smooth Workflow v4.0 - 03/14/2026
"WAN 2.2 Smooth Workflow v4.0" update: A new version with a new Ksampler method and more functions. Make sure to fully update your Comfyui and the extensions.
NAG: Now uses Normalized Attention Guidance (NAG) in Ksamplers. Check the github page more info. It makes the negative prompt be effective even at CFG 1.0 for improved quality and control.
Adaptive Prompts: You can use Adaptive Prompts when writing your Positive prompt now. Check the github page more info. To check the choices made by the Node just open the Positive Prompt Subgraph.
Updated Notes: Added hyperlinks for all Extensions used in the Workflow. Other notes have been added and updated. There are alot of little notes everywhere. Make sure to read them. ^^
The videos on Showcase used the SmoothMix I2V v.2 and SmoothMix T2v V.3.
All Extensions for WAN 2.2 Smooth Workflow v4.0:
Update: WAN 2.2 Smooth Workflow v3.0 - 19/01/26
"WAN 2.2 Smooth Workflow v3.0" update: Its a big workflow now. Make sure to fully update your Comfyui and the extensions.
All in one: Txt2vid, Img2vid and Audio Options! You can now try adding audio to your videos.
Select your workflow: I tried to make it very simple to change between the Img2vid, Txt2vid and Audio Workflows.
New fps Output: I had to change the video outputs to 25fps to make sure the audio and video are in sync.
Multiple Notes: There are alot of little notes everywhere. Make sure to read them. ^^
Added Extensions:
It uses MMAudio. So install - ComfyUI-MMAudio Nightly
And its requirements:
MMAudio NSFW Model (fine-tuned off the base model)
Puts the MMAudio models on ComfyUI/models/mmaudio. Create the folder if it does not exist.
Small Update - 10/16
Small change made to "WAN 2.2 S. Workflow v2.0" and "WAN 2.2 Txt2Video Workflow v2.0" That should allow for a faster VAE Decoding.
Update: WAN 2.2 Txt2Video Workflow v2.0 - 10/08
"WAN 2.2 Txt2Video Workflow v2.0" update: now its SUPER simple and easy to use.
I made it to use with the Smooth Mix Txt2Vid Wan 2.2 Checkpoint! Make sure to try it! ^^
Here are my recommend video setting when using the workflow with SmoothMix Txt2Vid:
Low/Mid end PCs:
Set Steps to "6" on both KSamplers;
KSampler 1 -> end_at_step = "3";
KSampler 2 -> start_at_end = "3".
High end PCs:
Set Steps to "10" on both KSamplers;
KSampler 1 -> end_at_step = "5";
KSampler 2 -> start_at_end = "5".
Update: WAN 2.2 Smooth Workflow v2.0 - 09/29
"WAN 2.2 S. Workflow v2.0" update: now its SUPER simple and easy to use.
I made it to use with the Smooth Mix Wan 2.2 Checkpoint! Make sure to try it! ^^
All the videos on the showcase for WAN 2.2 S. Workflow v2.0 used Smooth Mix Wan 2.2 Checkpoint.
It does not has Florence anymore. Make sure to install and update all necessary extensions listed below. No new extensions where used.
Update: Adv. Workflow - 08/29
"Adv. WAN 2.2 S. Workflow" update: added the Node "Color Match" before 'Video Result' and 'Last Frame' to help with color consistency.
Update: Txt2Vid Workflow
Added a txt2vid workflow. Thanks for vinque for the suggestion. =)
Update: Adv. Workflow + First2Last Frame
Made a few adjustments to the basic workflow for these new version!
Both use the same tools/weights as the basic workflow so no need to install anything new if you can already use the basic one.
Both versions have "Power Loader Node" set for the High and Low Loras.
More Loras usually means more instructions put on the Positive Prompt so Florence2 is disabled by default.
To make a looping animation on First2Last Workflow make sure to set enough time for the AI to process the loop. Try 5 or 6 seconds.
WAN 2.2 SMOOTH WORKFLOW
PLEASE make sure to read everything.
Making a simple Workflow from Wan is tricky. Wan requires are more complicated set up to work. I'm going into this assuming you have read all tutorials necessary to make Wan 2.2 work on your PC as doing that here is unresonable.
WAB 2.2 is still pretty new and there is still alot to test. All images used on the Showcase use the same set of loras, weights and configurations. The only difference was the image used and the Positive Prompt. I tried using realistic, semirealistic and anime style to show that its possible to animated any style you want by using the configuration on this workflow.
So, before we begin, this Worflow uses the following:
Use the one that fits your own specs. The one used for the showcase was the "wan2.1-i2v-14b-720p-Q8_0" version.
Out of all Loras that increases generation speed this one has given me more consistenty results.
The one used for the showcase was the "Wan21_T2V_14B_lightx2v_cfg_step_distill_lora_rank32" version.
Of course you should test other versions or Loras that speed up generation as well as well.
Showout to @CubeyAI for the Loras. Without these making NSFW results was a pain, and we can't have that can we? Make sure to use the latest version of those Loras.
Used extensions:
==========================================================
Its divided in sections to make it easy to navigate.
PROMPTS
Here is where the magic happens. So lets start here.
There are a few ways to use it:
TYPE NOTHING ON THE POSITIVE
By default Florence will create a prompt for you and a lot of times... thats enough lol. You don't need to type anything on the Positive. For examples of results that use an empty positive prompt check this and this.
I recommend you start by leaving the Positive Prompt empty at first. After you see the results you should begin working on your prompt if you did not like the results.
USE SIMPLE PROMPTS
Don't put a giant wall of text on the Positive Prompt as Florence will do that already. The showcase has Videos with very few prompts, some with just ONE.
DISABLE FLORENCE
If you really can't seem to get what you want disabling Florence might be the way to go.
FLORENCE
Choose your Florence model. Default model is "Florence-2-base-PromptGenv.2.0". Feel free to test other version or disable all of Florence nodes if you think its not helping.
CHECKPOINTS GGUF / SAFETENSOR
Very straight foward. Just choose your Checkpoint: GGUF or Safetensors. Disable the ones you are not going to use. Set to GGUF by default.
SPEEDING TOOLS AND LORAS
Here is where Sage and the Loras are set up. Feel free to change values or disable what you don't what to use.
CLIP / VAE
Still using the same as the wan 2.1
Text encoder: umt5_xxl_fp8_e4m3fn_scaled
VAE: Wan 2.1
IMAGE
Load your Image! :)
VIDEO SIZE
By default its set to 480x720 with Upscale by 2. That results in a 960x1440 video! Default upscale method is "lanczos"!
Feel free to change anything it to fit your own specs.
VIDEO LENGTH
Here you can adjust the video length.
By default is set to 81 or 5 seconds.
VIDEO SETTINGS
All settings here were set after testing. Feel free to change to your liking! :)
VIDEO PREVIEW/RESULT/LAST FRAME
You should be able to see how your video is turning out on the preview. If you don't where the results are going just cancel the run!
The video result is set to be at 30 fps. The video will be saved on Video/(Date_video_was_made).
The last frame of the video will be saved on the same folder as the video. You can use this to extent your videos! Check this post for examples.
====================================
Other recommend extensions:
Description
Added F2LF.
FAQ
Comments (136)
Just wanted to let you know that the description on this page mentions the "MMAudio SFW base model" but it links to the wrong repo. It should link to https://huggingface.co/Kijai/MMAudio_safetensors/tree/main (which you use in the workflow) instead of the repo from hkchengrex (which is not a safetensor and a LOT bigger).
Oh sorry about that. I put the correct link now. I really need to sleep. X_X
The NAG node cannot be installed in the new version.
are you trying to install from within comfyui or are you git cloning the version of NAG node he posted in the nodes used part of the changelog into the custom nodes folder, if trying to just get the extension manager to try install it try cloning it into custom nodes
@satsukiyami cloning it worked, thanks!
Love the workflow but it has an error, the VAE decode keeps erroring, I've updated all my nodes, comfyui and I've looked over the workflow, something between the second sampler and VAE decode is mismatched, I can't trace it though, anyone else have this issue?
I also get a VAE Decode error. I'm trying to use T2V w/ GGUF and I keep getting " TypeError: 'GGUFModelPatcher' object is not subscriptable "
@v_boissonneault282 Yeah, I've tried it on three ComfyUI setups, Desktop, Portable and Stable Matrix and it just doesn't work, I even swapped out the decode with the tiled version for stability and still, nope, not too sure what's causing the issue but I'll keep looking!
@v_boissonneault282 I've found a possible solution, run the text to video workflow, something simple, then turn it off and run the first frame and last frame, again, keep it simple, then turn it off and then run the image to video, I think the problem is tied to the VAE Decode needing something and for some reason the other two activate the search and install, try it, because mine works now!
@VicObXer thanks I'll try that right away!
So running the other workflows didn't work for me so I tried running it with LTX2 GGUF and it seems to work now. I could not get SmoothMix GGUF to work.
V5 is excellent, how can i speed it up a bit though with sage? how to attached sage attention nodes?
@DigitalPastel
One easy way would be changing the "Load Diffusion Model - HIGH/LOW" nodes for the "Diffusion Model Loader KJ" nodes. In those you can enable the sage attention.
F2LF works perfectly, I said I'd celebrate it and what can I say... I even put some beer in the fridge. 🥳🪇🎉🥳
HI i tried it but it didn't get the image from the f2lf at all. it's just random. how did you make it work
I'm having a problem with VIDEO Width x Height; I have a video box but I can't see anything. What could be causing this? https://ibb.co/XgTN89R
Oh! For that you disable "Modern Node Design (Nodes 2.0)" on the settings.
Ah ok thx
For all those who are having issues with the error "object is not subscriptable" I have found a solution, not too sure if it'll work for everyone but, if you run the text to video workflow first, then close it, then run the first frame to last frame, then close, then the image to video, it seems to fix any issues, for me, the VAE decode node was in a inactive state and refused to continue the process, the other two seem to force it to wake up and that resolved the issue, feels good to solve problems but again, I think that it may be the update for ComfyUI that caused the problem, hopefully everyone can start enjoying this workflow!
I tried what you suggested but it didn't work for me for getting I2V going. Instead, deleting the 'Clean VRAM Used' node just before the VAE Decode node got around the issue for me.
@eterelmel its work for me, thx dude
learning a ton from your workflows, thanks!
Installation Error: Installation failed: ComfyUI-NAG@unknown
have no idea how to fix shit
go to /custom_nodes in terminal, then clone this repo:
> git clone https://github.com/ThanaritKanjanametawatAU/ComfyUI-NAG-Fix
@Cyche Allelujah!
This works with AMD GPU? I got a Radeon 9060 and cannot do nothing, nothing works
Hi, it generates high quality videos. But it really is totally different from the input image, how can I get the start of the video as close as possible?
me too! and i have no idea how to resolve it.
could raising cfg help?
me too
Doesn't work. Says Enable TEXT2VIDEO for all toggle buttons instead of being able to choose IMAGE2VIDEO etc?
Managed to get it working by disabling the modern node boxes
I am getting the exact same problem...
I am getting the exact same problem...
I am getting the same problem and I have no idea how to fix it, I am very new to this.
@valfuze14390 Hello, I encountered the same problem. How did you solve it? I'm a newbie. Could you please tell me how to "disabling the modern node boxes"? Thank you very much
Please explain to me why creators of otherwise perfectly good workflows often use nodes that are impossible to deploy (for example, they don't pass site validation or are simply deployed with errors). This causes an incredible headache for inexperienced people like me, even though there are actually perfectly functional, widely used nodes.
Thank you! First 2 Last works like a charm! 💖
What GPU do you have?
Workflow doesn't work thanks to an uninstallable node named KSamplerWithNAG (Advanced). Why do people always do this? It's not like ComfyUI already has HUNDREDS of nodes, let's put some custom nodes impossible to install to have some fun.
Install this node package: https://github.com/ThanaritKanjanametawatAU/ComfyUI-NAG-Fix
worked fine for me
@Cyche Thank you!
For those of you struggling to get this working.... Re-install comfyui. I've found the portable version most successful. and no install! just extract. Easy. Then download the custom nodes via the links provided, extract them to your custom nodes folder in comfyui directory. Download models, and that's it. It works.
if you are running windows, I would use the official comfyui application.
The start and end images, for example, are set to last 8 seconds, but the final image always fades out in the last second, resulting in an uneven transition. Does anyone know why?
PLEASE REPLY ASAP hey i keep on getting this error whenever i try to use any mmaudio how do i fix this?? error: TypeError: BigVGAN._from_pretrained() missing 2 required keyword-only arguments: 'proxies' and
i have the same probleme
same problem here...
@Krankheit @marcelinorafik01419
Navigate to this path: ComfyUI/custom_nodes/comfyui-mmaudio/mmaudio/ext/bigvgan_v2/bigvgan.py
Next, change two lines 368/369:
proxies: Optional[Dict] = None,
resume_download: Optional[bool] = True,
at least this helped me.
Hello. I'll list my hardware first: I have an RTX 2070 8GB and 16GB RAM. The settings I used for the generation were: 480/520 resolution, 3 FPS, 6 steps, and I left the rest unchanged. Workflow - WAN 2.2 Smooth Workflow v5.0, Model: I2V v2.0 H/L, CLIP: umt5_xxl_fp16.safetensors, VAE: wan_2.1_vae.safetensors. I'll describe the problem: during generation, both KSamplers take about 3 minutes to complete the process. Then, when the process reaches VAE Decode, ComfyUI crashes without reporting any errors. At the end, it only says:
Windows fatal exception: access violation
Stack (most recent call first):
File "D:\ComfyUITest\ComfyUI\custom_nodes\ComfyUI-VFI\rife\rife_comfyui_wrapper.py", line 138 in interpolate_frames
File "D:\ComfyUITest\ComfyUI\custom_nodes\ComfyUI-VFI\nodes.py", line 140 in interpolate
and etc.
Does anyone understand this or have encountered this? Is there a way to get this process working? Maybe changing something in the parameters. I'd really appreciate your answers. Perhaps the problem is simply a lack of VRAM and RAM, meaning I need more powerful hardware. If you don't think it's worth wasting time on this because PC isn't powerful enough, you could let me know. I won't waste time researching it then)
The crash isn't in VAE Decode itself — it's in the RIFE frame interpolation node that runs right after. The access violation is almost certainly an OOM (out of memory) VRAM crash on your 8GB RTX 2070.
Quick fixes to try:
Disable or bypass the VFI/RIFE node entirely — see if the base generation works first.
If you need interpolation, set RIFE to use CPU instead of GPU in the node settings.
Lower resolution to 320/480 or reduce frame count.
Enable --lowvram or --novram in ComfyUI launch args.
Close everything else before generating.
Your hardware can run this, but it's tight. WAN 2.2 I2V + VAE + RIFE stacked is asking a lot of 8GB VRAM. Try disabling RIFE first — that's the most likely fix.
- by Claude
i am geting TypeError: 'ModelPatcherDynamic' object is not subscriptable pls help
Same. Dug around for about 3hrs before giving up. Hoping to come back to a fix someday
@zxcasddffdsxcv thank you so much!!!!
@zxcasddffdsxcv Thank you so much! I tried researching it myself and asking AI, but I couldn't figure it out all afternoon. You're amazing!
@zxcasddffdsxcv Bro, you are my Messiah
@zxcasddffdsxcv 10000 thumbs up!
Using this workflow, the generated videos, whether I2V or T2V, are always blurry.
Same here, i don't understand the problem...
Have you add the Loras ?
Try using https://civitai.red/models/1995784/smooth-mix-wan-22-14b-i2vt2v?modelVersionId=2513182 and updated loras, fixed this for me
@noisy_star yes 2214bi2v_i2vV20 step=30 cfg=2.3
Im using first frame last frame but it's not using my image at all when i generated it.
it's not getting the input of my image at all
@kenjanjan Same for me, there is an issue with the workflow somehow :/
Same here and I reinstalled "ComfyUI Windows portable" to be sure.
@noisy_star I think something changed in the nodes / Comfyui, cause when I used older version yesterday everything worked but now? Even older workflow on older model does not work :/
So i made a mistake, I downloaded the T2V version of low/high noise instead of I2V that's why it didn't took the input. @WwolfXX @noisy_star
Same mistake (names are confusing) I used T2V checkpoints.
Thanks @kenjanjan !
@noisy_star no probs bro!
@kenjanjan Could you tell me the exact name of the checkpoint I should download then? Cause im confused xD
@WwolfXX when you open up the sample outputs above. you can see the link of the checkpoint it used. it should redirect you to the model, then download the other variant.
i was facing blur problem any help, please 🥲
check the models if they are merged with or without LightX2V loras, blur is usually a lighting lora or ksmapler settings issue.
@Origion_0101 I don't have any lighting loras on and I'm using SmoothMixWan2214BI2V_i2vV20Low (and high). Should I be using something else?
This is a fantastic workflow - very easy to use, excellent tweakability/customization, extremely stable. Great work man!
Problem with promt generator:
[ERROR] LM Studio API error: Invalid URL '': No scheme supplied. Perhaps you meant https://?
Some one know how to fix thix?
im currently struggling on making the videos not blurry as heck when i use I2V. im newer to this stuff and im not exactly sure what the problem might be.
you probably need the light2xv loras:
- https://huggingface.co/lightx2v/Wan2.2-Distill-Loras/resolve/main/wan2.2_i2v_A14b_high_noise_lora_rank64_lightx2v_4step_1022.safetensors
Can it be that you missing are custom note you using because rife interpolation note is are red x ther and that means it is not ther anymore or not installed the only stuff that got this is the music node pack but its not ther any more maybe its now another name.Because i installed all custom nodes from you list.
"This workflow has an issue. After installing all the nodes listed by the author, a VAE decode error occurs. I switched to the wan2.2 sampler and the error went away, but the results are not as good as the author's. The problem seems to be with the New Subgraph node — I suspect the author did not package this custom node properly."
Unknown (4)
mmaudio_vae_44k_fp16.safetensors (1)
mmaudio_synchformer_fp16.safetensors (1)
apple_DFN5B-CLIP-ViT-H-14-384_fp16.safetensors (1)
mmaudio_large_44k_v2_fp16.safetensors (1)
Can you tell me which file I should download? And which folder should I put it in? Please.
Apperently you need to create folder yourself. like this C:\ComfyUI2\models\mmaudio
@Paddington27 tks
I asked chatgpt for help, and it was exactly as you said. I've now managed to create the animation. During installation, due to a version conflict, I had to use Stability Matrix and select ComfyUI v.019.3.
And I also had to follow chatgpt's instructions.
@hoabinhfsodfikj Yeah hahaha we dont understand AI so lets ask the AI how it wants it. xD
Been getting this error and cant' find a way to fix it The workflow does not contain any output nodes (e.g. Save Image, Preview Image) to produce a result.
Have you enabled the Workflow at the top in the Workflow Selection?
im a bit lost trying to find the SmoothMix_I2V_v2_High.safetensors (and low) models...... when I search for it it always gives me a v20 or smth. which outputs a blurry mess
Did you fix this? Becuase I have the same issue
@Piraterage23 issue persists
I have figured it out. Yuo need to use loras that linked in description, with this loras everything works fine.
@borod333 I tried with those before even making this comment, still blurry. Out of curiosity are you NVidia or AMD?
Excuse me I need help. It has an error when I try to make i2v. Did I do something wrong? Thank you.
MMAudio FeatureUtilsLoader
TypeError: BigVGAN._from_pretrained() missing 2 required keyword-only arguments: 'proxies' and 'resume_download'
In case you get an error when comfyui is importing "ComfyUI-NAG", you should look this :
https://github.com/ChenDarYen/ComfyUI-NAG/issues/53#issue-3629676396
What it says is that you have to modify some code lines in 2 files of the node pack.
Somewhere on your disk, you should have a folder for comfyui named "custom_nodes". You want to find those two files :
- custom_nodes\ComfyUI-NAG-main\chroma\layers.py
-custom_nodes\ComfyUI-NAG-main\chroma\model.py
You want to open those two files with a text editor and modify 1 import in each, save the files and restart comfyui.
From : from comfy.ldm.chroma.layers import DoubleStreamBlock, SingleStreamBlock
to : from comfy.ldm.flux.layers import DoubleStreamBlock, SingleStreamBlock
i wrote a tutorial on how to get this workflow running on SwarmUI for people who are new to workflows. its very casual and not meant at all to be an expert guide. i just wanted to save people (and future me) time troubleshooting! https://civitai.red/articles/31433/local-genning-tutorial-3-how-to-install-a-workflow-to-run-image2video-this-one-has-porn-in-it
goated, thanks for the hard work putting this together. just got it up and running, early results are spectacular
Regarding Smooth Workflow v5.0 AIO, in the I2V section, when I try to add a positive prompt, it appears as disconnected and I cannot add any text.
That's because of the absense of extension "ComfyUI-Adaptive Prompts".Install it and problem can be solved.
Pretty sure all the mmaudio stuff is now broken: TypeError: BigVGAN._from_pretrained() missing 2 required keyword-only arguments: 'proxies' and 'resume_download'
same here
I had that problem at first, too. I spent hours trying to figure it out and finally found the solution: Download the new version of mmaudio, or just copy the fixed bigvgan. py file into your mmaudio folder. It’s located under Custom Nodes -> mmaudio -> mmaudio -> ext -> bigvgan_v2. There’s already a bigvgan.py file there—just overwrite it with the new one, and it should work.
@LFPerfect is this with the NSFW model?
I don't think it matters which MMAudio model you use.
is it impossible to adjust the playback speed in Smooth? I increased both the frame rate (FPS) and the total number of frames, but the character's movements ended up severely sped up...
If anyone is experiencing this issue with mmaudio: Two required keyword arguments are missing from BigVGAN._from_pretrained(): “proxies” and “resume_download” ---> Solution: Download the new version of mmaudio or simply copy the fixed “bigvgan.py” file into your mmaudio folder. It’s located under “Custom Nodes” -> “mmaudio” -> “mmaudio” -> “ext” -> “bigvgan_v2.” There’s already a “bigvgan.py” file there—just overwrite it with the new one, and it should work.
Hello i have that issue ,i did action you mentioned and thanks but i have always MMAudio FeatureUtilsLoader
TypeError: BigVGAN._from_pretrained() missing 2 required keyword-only arguments: 'proxies' and 'resume_download'. And if i re copy en entire folder for bigvgan v2 comfuy lose completly the installation of mmaudio and impossible to install a bigvan version.
@RobJhonson Yeah, I had that problem with my new ComfyUI installation too, because I wanted a clean folder. I then shared the error with ChatGPT/Gemini, and it told me to add the following to “custom nodes->comfyui-mmaudio->_init_.py”:
import sys
import os
sys.path.append(os.path.dirname(os.path.abspath(__file__)))
from .nodes import NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS
Just open "_ini_.py" in your editor and add it like this, then save it, and it should work—it worked for me, anyway.
I always use an AI—usually Gemini—to troubleshoot.
@LFPerfect thanks => in _init_.py i have now :
from .nodes import NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS __all__ = ["NODE_CLASS_MAPPINGS", "NODE_DISPLAY_NAME_MAPPINGS"] import sys import os sys.path.append(os.path.dirname(os.path.abspath(__file__))) from .nodes import NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS
but i have always the same issue for the audio
i reloaded the workflow and i restarted completly comfuy
BigVGAN._from_pretrained() missing 2 required keyword-only arguments: 'proxies' and 'resume_download'
My version of comfuy version 0.27.1 ( the last version)
this is that init ? C:\Users\xxxx\ComfyUI-Installs\ComfyUI (2)\ComfyUI\custom_nodes\comfyui-mmaudio\mmaudio or C:\Users\XXXX\ComfyUI-Installs\ComfyUI (2)\ComfyUI\custom_nodes\comfyui-mmaudio ?
there is different version : GitHub - kijai/ComfyUI-MMAudio · GitHub
and GitHub - hkchengrex/MMAudio · GitHub
and if we don't install nodes by "nodes extension management" comfuy don't find the mmaudio installed
@RobJhonson Install mmaudio as usual via the ComfyUI Manager so that ComfyUI recognizes mmaudio.
Then copy the fixed version of bigvgan.py into the specified folder.
If you restart ComfyUI now, ComfyUI will no longer recognize mmaudio—at least that’s what happened to me.
That’s why I asked Gemini/ChatGPT what the problem was.
Gemini then told me that something was wrong with the paths or something like that—I don’t remember exactly what—and told me to do the following:
import sys
import os
sys.path.append(os.path.dirname(os.path.abspath(__file__)))
from .nodes import NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS
Enter it exactly as it’s written here into init.py. If there’s already something in the file—which will likely be the case—delete it and replace it with this, then don’t forget to click “Save.” I restarted comfyui afterward, and then it worked for me.
The issue is on the node MMAudio FeatureUtilsLoader ,What should these knots be pointing at? because bgvigan is not a safetensors
@RobJhonson Copy your error log into ChatGPT or Gemini.
@LFPerfect Okay, I started the whole process over from scratch and it's working now—thanks again for your patience.
Now I just need to figure out the audio part and where exactly to place the prompts.
Okay, the sound works, but what is it used for? Does it generate only noises or ambient sounds? Because no matter what prompts I use, I always get sounds—not human voices or sounds a human could produce.Could someone shed some light on this for me?
I asked ChatGPT, and that is indeed what it told me; could you confirm that this sound model is limited to this specific use?Thanks
@RobJhonson No, the MMAudio node (and the underlying MMAudio architecture) is primarily designed for generating sound effects (SFX) and atmospheric ambient sounds, not for human speech.
Here is a detailed breakdown:
Focus on SFX: MMAudio has been trained to generate sounds such as the clinking of glass, the rustling of wind, footsteps, or mechanical sounds based on an input video or text.
Lack of voice modeling: The model does not understand phonemes or human speech structure. If you try to force it to generate speech, it will at best produce vague, inarticulate sounds or noise that may vaguely resemble human breathing or mumbling, but never intelligible speech.
Different architecture required: For human voices (speech synthesis/TTS or voice cloning), architectures such as XTTS, Bark, Piper, or ElevenLabs APIs are used. These are based on entirely different training data (speech transcriptions) and models (Transformers for text-to-speech).
Conclusion: If you need a human voice for your project, MMAudio is not the right tool. Instead, you should add a dedicated TTS node (e.g., for Bark or XTTS) to your ComfyUI workflow to generate the audio component separately.
No, the MMAudio node (and the underlying MMAudio architecture) is primarily designed for generating sound effects (SFX) and atmospheric ambient sounds, not for human speech.
Here is a detailed breakdown:
Focus on SFX: MMAudio has been trained to generate sounds such as the clinking of glass, the rustling of wind, footsteps, or mechanical sounds based on an input video or text.
Lack of voice modeling: The model does not understand phonemes or human speech structure. If you try to force it to generate speech, it will at best produce vague, inarticulate sounds or noise that may vaguely resemble human breathing or mumbling, but never intelligible speech.
Different architecture required: For human voices (speech synthesis/TTS or voice cloning), architectures such as XTTS, Bark, Piper, or ElevenLabs APIs are used. These are based on entirely different training data (speech transcriptions) and models (Transformers for text-to-speech).
Conclusion: If you need a human voice for your project, MMAudio is not the right tool. Instead, you should add a dedicated TTS node (e.g., for Bark or XTTS) to your ComfyUI workflow to generate the audio component separately.
@LFPerfect Okay, thanks for the detailed answer and the confirmation—either use this tool for ambient sounds or turn to other models and workflows that require LTX-based models.I completely understand the need to use multiple tools. But after generating the audio, you still have to handle lip-syncing.
LTX models can do this now, but they still lack a lot of expressiveness.
Check out this account: https://civitai.red/user/Acaserio
@RobJhonson Yeah, I've seen accounts like that before, and they look breathtaking. But I haven’t really looked into it yet, since I’m still in the process of finding my own visual style.
But after a few minutes of searching, I found this:
https://civitai.red/models/2498991/dasiwa-ltx23-workflows-or-i2v-or-flf2v-or-t2v-or-audio
Why don’t you give it a try? The sample videos look really good with the synchronization, too.
I’ll definitely test it out when I’m ready.
@LFPerfect I’ve already tried an LTX workflow with audio; it worked, but it resulted in a singing voice (haha), even though nothing of the sort was specified in the audio prompt.
In my humble opinion, the solution is to clearly specify the action in the base video prompt for both MMAUDIO and LTX.
I ran a test using WAN video and MMAUDIO without a prompt for the MMAUDIO node, and it generated audio that actually matched the "base" prompt better.
I’ll be running more audio tests.
Of course, as you pointed out, the LTX process and audio results are more refined.
And the overall process is much more resource-intensive than WAN's.
In the output, I always get a blurry image that's different from the one I have at the input, but I don't have any video.
I run the workflows correctly. All the nodes except for the audio ones are installed and found.
But in the output, I don't have any video.
I'm working with an RTX 5070 (Comfy takes about 4 minutes to run, so it takes a long time for one image).
Thanks in advance for your help.
I was having the same issue but installing the Lightx2v Loras mentioned early in the post fixed the issue
@Darsovin thanks.i loaded 2 loras lightx2v mentioned in Power Lora Loader (rgthree) nodes ,there are enable.But i have the same problem i have always the last frame in png ,so no video :(
@RobJhonson try anothor checkpoint with lightning included like this: https://civitai.red/models/2003153/wan22-remix-t2vandi2v
@LFPerfect thanks i will try also that, but for the moment is select Wan2_1_VAE_fp32 and i have a VIDEO in OUTPUT.
Hello everyone, when i try to use First to Last frame workflow, my video gets some random highlight when playing back. Almost like the brightness bumps up suddenly. Any ideas why??
This workflow worked for me for the animation part only.
I couldn't get the audio working no matter what I tried. Looks like a recent update broke audio generation compatibility.
I also had to manually install "ComfyUI-NAG-Fix-main" (found it online) into custom_nodes because the NAG package wouldn't install through the Manager on my setup, and the old one kept throwing errors.
Finally, I was getting another compatibility error that killed the job. Fixed it by removing the "flow2-wan-video" folder from custom_nodes. Found that fix online too.
I had the same issue with the NAG node. I cannot get it to create anything more than a brown square no matter what length or size I make it though. And I can't get it to recognize the Clean VRAM as a latent output to save my life. Have you had a similar issue?
@Sweetpotatopie Nope, I didn't have that issue. But glad to see you figured it out! 😉
To anyone struggling to get past the TypeError: 'ModelPatcherDynamic' from the VAE, make sure you are using the workflow selector at the top center of the whole flow. For some reason, right clicking and setting a workflow group to "always" doesn't activate everything. Even bypassing the Clean Vram node just gave me brown images until I tried that.
Big thanks to @sfs123joel for solving it, and @demonicslop for making the tutorial page where I found the solution.
"Why does the Image-to-Video (I2V) workflow fail when I prompt a female character to undress? Instead of taking off her clothes, the AI usually ends up adding more layers of clothing.
Hello
Can someone explain to me how the resolution works?
Because the video I get as output is squashed with black bars on the sides.
Resolution (9:16 = 768x1376)
I have to set up nodes :
resizeimagev2 and wanimage to video
But even when tweaking these settings, I still get an output video that is squashed. And if it's a square image, the video frames are stretched in height.
Does the workflow keep the resolution 480x720?
Other things the face is totally very blurry and false ,same things with approprated and adapted negative prompt.
Excellent work for that complete workflow ,that deserves all our attention
ok i just select "resize" In the node resizeimageV2 and the video is cropped to the scale of original picture that loaded in input
v5 workflow is best overall!
But an older F2LF ran fine. But new one keeps randomizing the first frame, it's always slightly different from the base image. Why is this happening? All my diff models and loras are I2V
Got it, need to delete ImageFromBatch node and first frame will be yours making merged videos smooth again. I don't know what's the purpose of it besides ruining smoothness of videos longer than 5s, but without it everything fine. Workflow still not exellent, but very very good job overall!
