[edit:
13.05.2026: Update version 4.4 (see version description).
Small fixes to get back fast generations.
Attention:
If you struggle with node conflicts or you get errors while running the workflow, please have a look at my short Trouble Shooting Guide note in the wokflow first. Most importent is to update all components sucsessfully! ]
Special thanks to:
@ArcleinSK for investigation and solving the FLF issue, as well as forcing the First-Mid-Last Frame option and last but not least for charing fantastic knowlage.
@boinobin730 for initialising, forcing and supporting this project in all kinds of matter, like providing links, running tests, sharing knowlage and inspiring diskussions.
@Urabewe for publishing the original, perfectly running 12 GB VRAM LTX-2.3 workflows mainly used here in this workflow.
Features:
Simple to use all-In-One LTX-2 workflow with options for:
Text to Video
Image to Video
First/Last Frame to Video
Fisrt/Mid/Last Frame to Video
Video to Video
Text + Audio to Video
Image + Audio to Video
First/Last Frame + Audio to Video
First/Mid/Last Frame + Audio to Video
easy switching between all options,
all steps highly automated: no manual frame or width/hight calculations necessary,
easy to set inputs by predefined sliders and aspeckt ratio inputs (no risk to set wrong frame counts or wrong width/hight values),
completely automated resizing and cropping (if necessary) of your input images/videos.
brilliant audio generation (speech/sound) with LTX-2.3.
LTX-2.3 specifications:
Workflow version v4.3 consistently follows the LTX-2.3 specifications for 16:9/9:16 aspect ratios, including automatic width/hight calculations, as well as automatic input image/video resizing/cropping.
In addition you can simply choose now any other aspect ratios according to your needs while still getting the right values calculated for width/hight and automatic image/video resize/crop.
Requirements:
GPU with 12 GB VRAM (some users reported they got it running with 8 GB too),
32 GB VRAM,
Swap file size: 64 - 128 GB.
Speed and video length:
Runs very fast: 5 second (1280 x 864) Video: < 10 minutes.
Generation of long high quality videos in one run possible: 10 - 20 seconds without any issues,
Testrun: 30 second video (1024 x 704) tooks around 40 minutes without any OOM errors. Longer videos might be possible, but not tested yet.
Important:
This workflow is intended for advanced comfyui users who know how to install and operate the system and are able to resolve basic system errors themselves, like as node conflicts, or general system issues.
About this workflow:
This workflow is mainly based on the fantastic LTX-2.3 workflows of @Urabewe.
As far as I know, those were the first workflows running LTX-2 with 12 GB VRAM. All credits goes to the original creator.
My job was only to combine and organise the different workflows in a simple to use all-in-one design.
Description
Small fixes to get back fast generations:
re-edit of the audio part back to the template workflows,
reduce preview_rate back to 8 instead of 24.
Special thanks to @Silicon_Mirage for giving the right hint with the wrong audio part.
FAQ
Comments (64)
sorry i message you. ok so ... error .. Node '[[P:01 Text to Video]]' has no class_type. The workflow may be corrupted or a custom node is missing.
Great workflow, thank you!
Just the spatial upscaler link isn't right, nothing hard to solve.
I wonder if you could add a case for the lip sync LORA TV2V (https://huggingface.co/Lightricks/LTX-2.3-22b-IC-LoRA-LipDub and workflow https://github.com/Lightricks/ComfyUI-LTXVideo/tree/master/example_workflows/2.3)
Nice workflow; would love to see the version number in the actual workflow so one can know if they need a new one
Many thanks to the author for version 4.4, which he improved based on feedback.
Version 4.4 is very good — 245 seconds for 1024x1536 and an 8-second video 👍👍👍
9700x 32GB DDR5 6000 RTX 5070TI 16GB
I swapped out the GGUF loader for a diffusion model loader to use an FP8 eros check point is the only real change I made; But I'm noticing that after every generation the next one will take just over an additional 100 seconds on top of the previous time
I'm using a 5070ti 16gb, 32gb system ram and a 64gb swap file. Any idea on ways to implement something that will prevent this kind of behavior? Thank you!
Edit: going to try the v4.4 workflow and see if that helps!
My generated videos always come with unwanted background music automatically, how can I turn it off completely?
maybe add what you want to hear in the prompt "quiet room ambience, soft footsteps, distant traffic, subtle environmental sounds, natural room tone, unscored scene, documentary style audio, no musical score". Or describe the kind of music you actually want to hear... Otherwise it will try to fill the blank with something (music most of the time... as in movies)
Great a new version.
I am looking forward to testing it ! Thanks Arkinson !
I'm having problems with the missing node switch for the input options. Where do i get that?
Not sure if you figured it out, and If you mean for the node in "Input" blue group, top left, it's https://github.com/tritant/ComfyUI_Custom_Switch
Version 4.4 definitely improves the generation time. At least for the few tests of batched generations done so far. Thank you Arkinson !
This workflow is one of the most straightforward and well documented I've used!
Is it possible to a V2V with a source video that has no sound? I'm getting an error "LTXV Audio VAE Encode Exception: VHS failed to extract audio from ... .mp4. If not, is there a way to have it create sound for the video first?
画像から映像/参照への機能で問題を抱えている人もいるのに気づきました。詳細な構成については私のウォークスルー動画で説明しています。
参考画像の説明は3分20秒から始まります。もし行き詰まったら、ぜひご覧ください!
I've done this with with a modified version of this workflow. See this post for video examples. The workflow is embedded in the images.
https://civitai.red/posts/28200204
@RandomAIUser its work! blank audio - https://github.com/anars/blank-audio/blob/master/1-second-of-silence.mp3
I keep getting really blurry and smeared videos, idk why, even if I keep the text prompts simple or complicated, it's the same (rtx 4070 ti 12gb)
You could be using a wrong model version. Try downloading the exact models by name that are noted in the workflow.
Download this file,
ltx-2.3-22b-distilled-lora-dynamic_fro09_avg_rank_105_bf16.safetensors
And put it in ComfyUI/models/Lora
I had the same issue as you, and that file was missing, but I didn't know.
Also, go to the Workflow on your computer,
and in the far left it says "Model Links".
Go look at each one, and make sure that it is downloaded and in the file location that it says at the bottom of that Model Links section.
I was missing the Destilled Lora, even though ComfyUI would still run... it would just make hyper blurry videos.
Also, tip, make sure that you download the files in the links at the top, because the file names at the bottom have some typos (some of them have the old names).
Let us know what happens!
The video generations are blurry, although fast.
Is this working for other people?
I apologize if this is stupid answer, but perhaps you're running low steps without distilled LoRA.
@Arcadeath I don't know!
I will check.
Thank you for giving me some insight, though.
@maninstrawhat You're welcome. You could always check if your model requires a distilled LoRA. I use Eros10 model and I use lltx-2.3-22b-distilled-lora-dynamic_fro09_avg_rank_105_bf16.safetensors for first pass, and ltx-2.3-22b-distilled-lora-384-1.1.safetensors for final pass.
@Arcadeath That's cool!
So... this workflow uses the ltx-2-3-22b-dev-Q4_K_M.gguf model, but you use the Eros10 model?
By the way, Arcadeath, do you know how I can check if I am using the right Lora?
Because I see the "Power Lora Loader (rgthree)" node, and it says that the toggled model is "90 video\ltx-2.3-22b-distilled-lora-dynamic_fro09_avg_rank_105_bf16.safetensors" with Strength of 0.60 , however I don't see that file in my ComfyUI folder?
So, I'm wondering if it's really loading, and where it might be at.
Also, in order to add another pass for the "final pass", could I just click the "+Add Lora" node button? And then add model, like what you are using, for instance?
Thank you for your help!
I really appreciate you taking the time to respond to me!
It makes a big difference,
Edit: I just downloaded the LoRA file in question, and put it in the Models\loras folder... so, here is to hoping that works. *shrug*
@maninstrawhat You're welcome! I think that should fix your issue, but I am not entirely sure if this workflow has first / final pass separated as the one I use (I actually want to try this workflow but didn't yet). Also, if you have 12GB VRAM + 32GB RAM you can use models better than gguf q4 (IF heh).
@maninstrawhat I just tested this workflow with and without LoRA! Without LoRA it's blurry/smeary, with LoRA it's perfect for me. Still, I don't really know what kind of LoRAs quantized models need (GGUF ones). What's your VRAM and RAM amount?
@Arcadeath I am running what it suggested: RTX 3060 12GB, and 32GB RAM.
Actually... Arcadeath, I got it working last night! :D
It seems to be what you suggested.
I think that I had the right lora selected in ComfyUI, but when I went to double check that it was actually installed, I didn't find it at all in the ComfyUI directory.
So, I downloaded it by doing a G**gle search, and put it in the models/loras folder , restarted ComfyUI... and it worked!
I was wondering... how would it even work at all without the right LoRA ?... this is all new to me.
I really appreciate your help!
Because I followed your suggestion, and it worked out!
I greatly appreciate it! :)
I'm glad that it works for you, too.
I guess next I'm going to try out adding another LoRA.
And a better GGUF model, too!
Thank you for the tips about using a better model. I want to try that now.
@maninstrawhat Hey, I'm happy it helped! So, what really happened and why it's working now: Models require lots of steps to generate a coherent output. If you increased amount of steps to 30-50 instead of 8, it would generate nice results probably too. But that would take 5 times longer sadly. So people make distilled LoRA which can reduce amount of steps (like turbo image models). It reduces quality by tiny bit, depending on models, but way way faster. Also GGUF models are way slower (quantized models), and they can't be trained upon further. My recommendation is to always get diffusion models since you can offload their size to RAM (hence the increasing RAM cost). Best way to know is if you check VRAM usage per system (I use Linux with 300MB usage), then add model GB, add LoRA, text encoder, clip, and all other files as size. If it is higher than your total VRAM + RAM then you can't technically run the model. This workflow is well optimized since it has great offload switching, model unloading when they are not used, etc, so you don't out of memory (OOM). fp8 models are bit faster on 40xx series because of native support, I got 4070ti, but they can be still ran on 3060.. Even then I think they will be faster for you than GGUF models. Also Q(number) GGUF models are ranked by size. Just because a model is 20-30 GB doesn't mean that you can't run it because you can offload parts of it to RAM. Some parts are not unloaded, some settings can be tweaked, but with 12GB VRAM and 32 GB RAM I can generate 2k videos of 15 duration with this workflow and Eros10 fp8 diffusion model. GGUF files = slow but lower size, I could run Q8 model too, but it's still worse than fp8 one in both terms of speed and quality at that point. Diffusion models need separate clip and vae files, and there are also checkpoint files which have vae and clip baked in, so you need one file instead of 3 (merged models). LoRAs are just 'addons' to existing models, like a new concept, and stacking them also increases memory usage since they are loaded with your model. So, the question is can you fit that into 48GB with system usage, apps open etc. I also have 64GB .swap file allocated on my SSD so it can use it as very slow reserve RAM.
@maninstrawhatAlso, to add different LoRAs to first and final passes just copy paste lora loader with those two tiny nodes connected to and connect model to both of it. Setter of lora loader 2 will automatically update. You can double click that main "Processing" node and on bottom left in new window you can make another setter. Processing Is actually like "folder" node to organize background stuff. Inside you can see how models are loaded for first and final pass. In this workflow they are called earlier, then they got fed into both first and final pass models at once. There you can feed new getter separately into final pass (Follow the lines and where they go and try to understand them, search for first get_m model. It's best ComfyUI learning experience. In this specific one when you're in processing, zoom out a lot, and long curvy line in top middle is model load for final pass. It goes down right into it (when that node is green during 3/3 processing in terminal). Click on nodes and you can open their documentation on the right. Some can be double clicked to expand since they work like back end organizers only).
I can send it after weekend once I got more time. You can see at very beginning how model's are "SET", you need to set another one to different name, and "GET" it separately into first and final pass, you need to reconnect the line with new getter into final pass node.
Follow terminal and what node is green is best way of understanding what happens when.
I just finished installing the workflow... Double Checked all of the model names they are all correct. I am on the latest portable version 0.22 any ideas?
Tried the v.43 workflow as well. getting the same results
how without ANY info whatsoever..
Do you have the Destilled LoRA in your ComfyUI/models/loras folder?
Because I found out that I did not, and after downloading it and manually inserting it into the folder, the workflow worked.
Before I did this it was generating very blurry videos.
This is the name of the file I was missing:
ltx-2.3-22b-distilled-lora-dynamic_fro09_avg_rank_105_bf16.safetensors
I would check on the far left side of this workflow, in the "Model Links" section, and make sure that you have every file it lists present inside of your ComfyUI install folder, and in the listed location.
It says at the bottom of that section where the file needs to go.
Be careful to use the model names at the top, though, because at the bottom it has some typos.
Hi, is there a possibility to add separate LoRAs for first 8 steps only, then a different LoRA for final 3 steps? I want to use different distilled versions for first and final steps.
NEVERMIND I just found how to do it, great workflow, amazing and simple to edit!
Is there any way to increase the steps? Small details are a bit blurry like eyes, fingers with motion for example. Best workflow I have used for LTX 2.3 so far. I don't mind sacrificing speed for the needed quality.
In each subgraph there are two nodes called ManualSigmas with a bunch of numbers in them. Each of those numbers is a step. The easiest way to change it so you have control of the number of steps it to replace the Sigmas node with a 'Basic Scheduler' node. Use 'linear_quadratic' as the scheduler and adjust the number of steps to whatever you want.
@RandomAIUser Thanks for the tips, it is improved after testing combinations. I added the LTXVScheduler node and use 16 steps on first sampler, and this sampler alone the quality of video surpass the final from the original workflow. Then on second sampler I used the Basic Scheduler(linear_quadratic) combined with the Sigma Rescale node(set to start at .75) for the 3 step.
@R3G4L Hey there. I see something interesting is happening here. Can you share your workflow?
@R3G4L I have had really good luck with setting the first stage to 2-3 steps and 8 steps for the second. I know that is totally backwards to what they are supposed to be. But the quality (and prompt adherence) after the second stage seems better for some reason. It almost looks like it starts over with the second pass. Making me wonder what a single stage workflow would look like. But I haven't tried it.
couldn't wait, so I set it up myself =). In the end, for the first pass, I set CFG to 3, 30 steps, denoising in the Base Scheduler node to 1, plus various LoRAs. For the second pass, I set the Distillation LoRA to strength 0.6, denoising in the Base Scheduler node to 0.45, and 9 steps. The face stays consistent without any hocus pocus in most cases. And the quality seems to have improved. I'm still testing it, but overall it's OK.
This is a great workflow. It works fantastically and is easy to understand, but am I the only one having issues with speech lining up with lips in the audio vs video? Like, if I have spoken dialog in a video, the audio will usually turn out fine, but the lips never match up with the speech. Maybe this is just a limitation with ltx 2.3 that I wasn't aware of, but wanted to check.
Much appreciated.
sometimes, but adding this helps:
her mouth opens and closes naturally with each syllable as she says "blah blah blah"
generally only need to say it the first time, but if there's a lot of movement between dialogs or multi characters speaking, sometimes you need to add it at each problem area.
how to edit video its not working moine
How to remove negative from this workflow?
The most stable and simple version of this workflow is the 2.0
I can confirm that version 4.4 seems bugged but it is most likely my fault for not understanding ComfyUI better. I went to version 2 where everything works perfectly fine.
I make one video 2 min with 480p.
Is there anyway we can do it faster?
Example : I don't need audio but it keep create. How to prevent it?
"Missing custom nodes (1) This workflow uses custom nodes that you have not yet installed.
comfyui-various (5). I have already installed it via git clone https://ghproxy.com/https://github.com/jamesWalker55/comfyui-various.git but it still shows this error. How can I fix this?"
did u find a solution to this?
I got rid of that missing custom node. One of these commands did it. Not sure which.
Python Package > + > "pip install" opencv-contrib-python
Python Packages > kornia package > downgrade it to the version below 8.0
install triton-windows
Package Commands > Install Bitsandbytes (ROCm)
I'm failing to install "comfyui_layerstyle" nodepack though. Got any tips how to install it?
Probably related to the opencv-... (it needs opencv-contrib-python) but comfy seems to auto-install other opencvs that conflict with it. Couldn't solve that one yet. Any ideas?
The issue is that ComfyUI uses its own virtual environment (.venv) with a different Python version, but soundfile was installed into your system Python. ComfyUI cannot see packages installed outside its virtual environment.
To fix this, open a command prompt and run the following command. Replace the path with the actual path to your ComfyUI installation:
"C:\Your\Path\To\ComfyUI\.venv\Scripts\python.exe" -m pip install soundfile
For example, if your ComfyUI is installed in D:\ComfyUI, the command would be:
"D:\ComfyUI\.venv\Scripts\python.exe" -m pip install soundfile
After running this command, restart ComfyUI Desktop completely. The missing nodes (JWInteger and JWIntegerToFloat) should then load correctly.
I'm failing to install "comfyui_layerstyle" nodepack.
Probably related to the opencv-... (it needs opencv-contrib-python) but comfy seems to auto-install other opencvs that conflict with it.
Couldn't solve that one yet. Can anyone help?
Same here
I've started experiencing errors during the video encoding process. It says Video Helper Suite fails because wrong values are given to ffmpeg, and it occurs only when I chose H.264 and H.265 as the outputting format.
Can anyone ADD to the current workflow THIS option? https://civitai.red/models/2757549/best-faceid?modelVersionId=3103007
If anyone is having the issue with comfyui-various. Install the git from: https://github.com/jamesWalker55/comfyui-various
Then in the comfyui root folder, open cmd and paste : .\python_embeded\python.exe -m pip install soundfile
(For Reference: https://github.com/jamesWalker55/comfyui-various/issues/27)
i manually replaced all comfy-various node to alternative, especially those float to int. conversion, comfy-core have similar function. Also can use those hot/common node like rgthree kjnode....
donn't why author use so many different custom node while some of them are not necessary...
Where can I download comfyui_custom_switch for this Workflow
@arkinson FYI, your WF is still my go to Wan 2.2 method. Have tried many others but yours still seems to be the best. Thanks.
Trash that causes comfyui to fail to load
Amazingly smooth workflow.
But I can't seem to disable exporting the video that doesn't have audio. Does anyone know of a way to do this with the VHS Combine Video node? We all know how valuable HDD space is here...
Doesn't work , the least a flow should do is produce an output that is not burry or completely deformed from default settings if someone gets the exact files you list - I'm so sick of waisting time trying to fix workflows that produce garbage