[edit:
13.05.2026: Update version 4.4 (see version description).
Small fixes to get back fast generations.
Attention:
If you struggle with node conflicts or you get errors while running the workflow, please have a look at my short Trouble Shooting Guide note in the wokflow first. Most importent is to update all components sucsessfully! ]
Special thanks to:
@ArcleinSK for investigation and solving the FLF issue, as well as forcing the First-Mid-Last Frame option and last but not least for charing fantastic knowlage.
@boinobin730 for initialising, forcing and supporting this project in all kinds of matter, like providing links, running tests, sharing knowlage and inspiring diskussions.
@Urabewe for publishing the original, perfectly running 12 GB VRAM LTX-2.3 workflows mainly used here in this workflow.
Features:
Simple to use all-In-One LTX-2 workflow with options for:
Text to Video
Image to Video
First/Last Frame to Video
Fisrt/Mid/Last Frame to Video
Video to Video
Text + Audio to Video
Image + Audio to Video
First/Last Frame + Audio to Video
First/Mid/Last Frame + Audio to Video
easy switching between all options,
all steps highly automated: no manual frame or width/hight calculations necessary,
easy to set inputs by predefined sliders and aspeckt ratio inputs (no risk to set wrong frame counts or wrong width/hight values),
completely automated resizing and cropping (if necessary) of your input images/videos.
brilliant audio generation (speech/sound) with LTX-2.3.
LTX-2.3 specifications:
Workflow version v4.3 consistently follows the LTX-2.3 specifications for 16:9/9:16 aspect ratios, including automatic width/hight calculations, as well as automatic input image/video resizing/cropping.
In addition you can simply choose now any other aspect ratios according to your needs while still getting the right values calculated for width/hight and automatic image/video resize/crop.
Requirements:
GPU with 12 GB VRAM (some users reported they got it running with 8 GB too),
32 GB VRAM,
Swap file size: 64 - 128 GB.
Speed and video length:
Runs very fast: 5 second (1280 x 864) Video: < 10 minutes.
Generation of long high quality videos in one run possible: 10 - 20 seconds without any issues,
Testrun: 30 second video (1024 x 704) tooks around 40 minutes without any OOM errors. Longer videos might be possible, but not tested yet.
Important:
This workflow is intended for advanced comfyui users who know how to install and operate the system and are able to resolve basic system errors themselves, like as node conflicts, or general system issues.
About this workflow:
This workflow is mainly based on the fantastic LTX-2.3 workflows of @Urabewe.
As far as I know, those were the first workflows running LTX-2 with 12 GB VRAM. All credits goes to the original creator.
My job was only to combine and organise the different workflows in a simple to use all-in-one design.
Description
Minor update after testing and several very usefull user inputs:
bug fix: Aspect Ratio subgraph: changed round_to_multiply = 64 insted of 32,
some little improvements, like:
bypassing audio for the upscale pass,
adding audio preview,
increasing preview_rate = 24 for better video previews.
FAQ
Comments (56)
Hi - First! Thank you for all your hard work!
I too hit the NaN/+-Inf [aost#0:1/aac @ 0x5cd2c08dadc0] Error submitting audio frame to the encoder after a Comfy update.
However, I'm still getting the NaN with the latest 4.3 ver of your workflow. (using the out of the box settings) I'm also now getting OOM unless I drop down to the 3Q gguf.
Background and possibly helpful info. I'm on Ubuntu. I have an RTX5060 16gb vram. 96gb system ram.
Before the ComfyUI update I was able to run the Q8 gguf (both dev and distilled) versions of the models without any issue and produced vids up to 10 secs.
I've updated everything via manager in ComfyUi. I still get the errors on 19.5 19.4 and 19.3 versions of Comfy. Certainly seems to be Comfy induced. I can still make vids with ltx2.3 on ltx desktop without issue.
@piehound0101723 "I still get the errors on 19.5 19.4 and 19.3 versions of Comfy"?
Sorry, but I`m really not sure what you are talking about. Latest comfyui version is 0.19.3. I am updated today and everything works as usual - pleas look here too.
@piehound0101723 Wich OS and comfyui version and release version do you really use? Anything broken during the update??
@arkinson I am on Ubuntu 24.04.4 OS
Comfy manager tells me it is 19.5 -> https://imgur.com/a/TyznBx2
Why that is different from the main branch I am not sure.
But some good news. Notice that updated KJNodes? After that I no longer get the NaN/+- error! (I made sure custom_scripts) was correct after I took that screen shot.
However, I am still getting OOM if I go to a model bigger than Q3 gguf. Which is odd because I was able to run Q8 gguf before. As I said previously I have a 5060 16gb vram.
@piehound0101723 Imgur do not open your screenshot, but on the Linux part I`m out - sorry.
@arkinson No worries - understand. I was not blaming your flow. Comfy has broken something in their updates. I should have mentioned in the previous update I was getting the NaN/+- error in @Urabewe's flows as well as yours. But at least the latest Comfy update fixed that. Now I just need to figure out what Comfy broke that is causing the OOMs.
@arkinson Sharing here in case someone else hits the issue. I had to update my Comfy startup to be
python main.py --reserve-vram 3.0 --lowvram --disable-pinned-memory
Note - you may be able to lower that --reserve-vram number. I posted a little more info on a thread in reddit on /comfyui -> https://www.reddit.com/r/comfyui/comments/1svix8a/oom_errors_after_comfy_update_and_how_im_getting/
In version 4.3, generation time has almost doubled: an 8-second video at 1536x864 now takes 562.48 seconds, compared to 300.90 seconds in version 4.2.
4.3's rendering time has increased critically: 3/3 [04:38 < 00:00, 92.89 seconds/it] compared to 3/3 [01:48 < 00:00, 36.31 seconds/it] for 4.2.
Also, in version 4.3, model initialization for the second pass may not start and takes a very long time. Refreshing the browser helps.
My PC is AMD 9700X 32GB + 4070TI 12GB.
@pavelinet87445 Please run a simple test:
Open main subgraph and go to "LTX2 Sampling Preview Override" node and set preview_rate = 8 instead of 24.
Let me know if this works for you.
@arkinson
Thanks for the quick response.
No, increasing or decreasing preview_rate = 8 or 24 doesn't affect generation time.
Furthermore, the model can't initialize on the second pass unless the video memory is completely cleared. Here's my log from the first generation, but on the second, everything freezes.
Generation 1 log:
100%|██████████████████████████████████████ ███████████████████████ ██████████████████████| 8/8 [02:19<00:00, 17.49s/it]
Requested to load AudioVAE
loaded completely; 693.46 MB loaded, full load: True
Unloaded partially: 2255.02 MB freed, 6229.20 MB remains loaded, 65.70 MB buffer reserved, lowvram patches: 404
Requested to load VideoVAE
Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.
loaded partially; 9806.84 MB usable, 9761.75 MB loaded, 4138.70 MB offloaded, 45.08 MB buffer reserved, lowvram patches: 0
100%|█████████████████████████████████████ ███████████████████████ ██████████████████████| 3/3 [03:01<00:00, 60.57s/it]
Requested to load VideoVAE
0 models unloaded.
Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.
Prompt executed in 396.45 seconds
2nd generation log:
got prompt
Requested to load LTXAV
0 models unloaded.
Unloaded partially: 1338.43 MB freed, 8423.32 MB remains loaded, 45.12 MB buffer reserved, lowvram patches: 1415
100%|█████████████████████████████████████ ████████████████████████ ███████████████████████| 8/8 [02:17<00:00, 17.15s/it]
Requested to load AudioVAE
loaded completely; 693.46 MB loaded, full load: True
Unloaded partially: 2246.69 MB freed, 6176.63 MB remains loaded, 65.70 MB buffer reserved, lowvram patches: 1807
Requested to load VideoVAE
Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.
loaded partially; 9806.84 MB usable, 9761.75 MB loaded, 4138.70 MB offloaded, 45.08 MB buffer reserved, lowvram patches: 0
Attempting to release mmap (1893)
0%| | 0/3 [00:00<?, ?it/s, Model Initializing ... ]<--- Here the model can't initialize; only a full restart of run_nvidia_gpu.bat helps.
I think the problem is in the memory clearing method change.
Log:
Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.
partially loaded; 9806.84 MB usable, 9761.75 MB loaded, 4138.70 MB offloaded, 45.08 MB buffer reserved, lowvram patches: 0
Attempting to release mmap (1893)
However, on version 4.2, even with preview_rate = 24, everything works. Here's the log from version 4.2:
100%|█████████████████████████████████████ ███████████████████████ ██████████████████████| 8/8 [02:16<00:00, 17.07s/it]
Unloaded partially: 1535.73 MB freed, 7188.46 MB remains loaded, 65.70 MB buffer reserved, lowvram patches: 309
Requested to load VideoVAE
Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.
0 models unloaded.
Unloaded partially: 2905.41 MB freed, 4283.05 MB remains loaded, 180.31 MB buffer reserved, lowvram patches: 587
100%|████████████████ ██████████████████████ ███████████████████████ ███████████████████████| 3/3 [01:50<00:00, 36.82s/it]
Requested to load AudioVAE
loaded completely; 693.46 MB loaded, full load: True
Requested to load VideoVAE
0 models unloaded.
Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.
Prompt executed in 322.01 seconds
@pavelinet87445 I also encountered this problem. I checked the changes in v4.3, and the author removed the LTXVConcatAVLatent node before the secondary amplification node. He directly connects the amplification node sampler via latent. I think this is the key to the problem. When I put LTXVConcatAVLatent first, it returned to normal speed. However, the audio latent must be connected to the LTXVConcatAVLatent node, otherwise it will report an error. But you can connect the audio VAE decoder before the first sampler to bypass the amplification sampler. In the separate audio and video latent space of the amplification sampler, you can leave the audio latent output empty.
@Silicon_Mirage I would be very grateful if you could share your workflow correction, because I have just started and cannot edit the workflow so deeply.
@pavelinet87445 Hi, and sorry for the late response. I had not much time over the last days. Thank you for testing the preview option and please see below.
@Silicon_Mirage You are right, bypassing the audio part over the second KSampler seems to cause the explicit longer generation times. Thank you so much for the hint 👍
I did a couple of tests meanwile and I will go back to the standard audio processing.
I also did some more serious speed tests and mostly it runs faster with the lower preview rate = 8.
I will try to publish a new workflow version soon. Please be patiant.
@Silicon_Mirage @pavelinet87445 Workflow version 4.4 is out now.
What do I need to do to solve this problem?
RuntimeError: Input type (struct c10::BFloat16) and bias type (struct c10::Half) should be the same
Thanks for the update.
I have no problems with your workflow. I tested both the new and the previous workflow. The speed is the same (I2V, length 10 seconds & LoRa only about 270 seconds). My setup: ComfyUI Easy Install, RTX 4070 Ti, ComfyUI 0.18.1, ComfyUI_frontend v1.42.10.
In my experience, you should use the workflow with a clean ComfyUI restart. The memory management is very sensitive.
Everything seems fine until its time to run the audio vae, and then the video turns into a psychedelic mess of colorful grids. Can anyone help?
Check the latent_upscale_model. ltx23 requires 2.3 model.
What settings do you recommend for my system ?
RTX 3090 24gb , Ryzen 9600X , 64 gb RAM
gen 4 SSD
and thanks for your hard work
Thank you for this excellent workflow. I dont really know what im doing but this made it much easier.
The videos turn out much better than the WAN ones i was doing before.
Hi, I am getting hard cuts and frame jumps whenever I use FF/LF (v4.3).
Looks like the LF is inserted way to early, I'm not sure.
The clearly visible workflow and easy-to-understand piping are truly amazing. ❤
Is there a way to use LTX2.3-10Eros in the workflow?
Hi @arkinson I'm glad to see you are still plugging away at your wonderful workflows. I have been away for over a month now due to unforseen work (not likely to end soon), Hopefully around end of June I might get to play with comfyui again. I wanted to check in on Civitai as well as how you are going these days. I hope you are doing well.
@boinobin730 Hi - I`m not really active here for myself actually. Finally I am not glad with the changes in the civitai system. Seems the fun times are passing by - slowly but surely....
@arkinson I got busy and only just realised they separated the SFW away from the NSFW. It feels like an aeon has passed since I was generating. My confyui is still .181 LOL. Too scared to update AS I don't have time to troubleshoot anymore. I will get time again. eventually.
It's a shame you have weaned yourself off Civitai. I will miss our chats about new workflows, models, checkpoints, LTX problems. Hopefully you will come back to it eventually. Look after yourself.
@boinobin730 I`m not away, but actually I have not the time and motivation to stay much active here.
Any idea why when running the very same workflow after it ended a rendering, it might take ages to complete the second one ?
it keeps indicating Estimated to finish ~1 sec (since 30 minutes). First run took only a few minutes... There must be something I don't quite get... Any help would be welcome :-) New WF is great by the way...
I don't get it, some of the video I posted disappeared from this page. As far as I know they didn't violate any civitai rules... Did this happen to other persons as well ? Any idea for a reason ?
Requested to load AudioVAE
Unloaded partially: 625.92 MB freed, 12903.52 MB remains loaded, 13.77 MB buffer reserved, lowvram patches: 279
loaded completely; 768.90 MB usable, 693.46 MB loaded, full load: True
Unloaded partially: 52.40 MB freed, 12851.12 MB remains loaded, 13.77 MB buffer reserved, lowvram patches: 298
loaded partially; 13682.35 MB usable, 13669.76 MB loaded, 259.99 MB offloaded, 12.59 MB buffer reserved, lowvram patches: 0
Well, hell! Seems comfyui doesn't want me to use this. Keeps telling me I am missing the LayerUtility: AnyRerouter no matter what I do. Any suggestions?
English Support: If you encounter the "AudioVAE... keyword argument 'sd'" or "Class Type" errors, please check the troubleshooting section at 02:48. I've listed the fix commands and code snippets in the description below!:https://youtu.be/kEBgAkVHlBE?si=9XKhF1-YfLbRB9cv
@sanmarinogmip394 I will check it. Thanks!
1:31から見てください
I eventually finally got it to work using claude cowork which actually appears to have successfully rewrote the workflow for me (can't upload it here, alas). There may be other things you need to do prior to this, but the below pointed out what finally got me through that issue. After that all I needed to do was apply the latest update of KJNodes to get the preview to work properly. Now I just need to figure out why it isn't using the negative prompt, but it's rendering video with audio now (albeit the audio is on a separate track and would need to be merged). edit had to increase the two CFGGuider values to 3, seems to adhere better to neg prompts now...
Claude:
Found the answer. Every LayerUtility: AnyRerouter is a simple 1-input / 1-output wildcard passthrough — input named "any", output named "any", both type *, no widgets. Functionally identical to ComfyUI's built-in Reroute node. All 19 of them live inside one big subgraph (a786c8d3).
So the cleanest fix is to rewrite the workflow to use Reroute, which is a built-in ComfyUI node with the same shape. Let me make you a fixed copy of the workflow so you don't have to swap 19 nodes by hand.
https://comfy.icu/node/LayerUtility-AnyRerouter I downloaded the .rar from the git link here and replaced the entire custom node folder: "ComfyUI_Layerstyle" with the "ComfyUI_LayerStyle-main" folder instead (do not keep both, only use the -main layerstyle folder) and it worked after. ig the installed extension for layerstyle on comfy didnt include the rerouter component internals.
Thanks guys, for all the help. I will get to trouble shooting this thing soon with the advice given.
https://youtu.be/kEBgAkVHlBE?si=KnFza8HMJxW2yZT2
English Support: If you encounter the "AudioVAE... keyword argument 'sd'" or "Class Type" errors, please check the troubleshooting section at 01:31. I've listed the fix commands and code snippets in the description below!
will you add Sulphur LTX 2.3 support?
You just need to replace the main model to use it—that's right, that's exactly what I did!
@Silicon_Mirage it has a new prompt tool
Thank you for this! Super easy to setup as someone who has very little experience in comfy and workflows and I just used the FP8 eros unet. Execution times are really good with my 5070ti 16gb and 32gb system ram.
Guys, I’m kind noob to this stuff, but I noticed that the LTX I2V, when I use First/Last Frame, changes the video output a bit. It’s not exactly the same as the last Frame reference image.
Something that doesn’t happen with WAN22 (at least for me). Is this normal behavior for the model, or am I doing something wrong? I was hoping the last frame would match the reference exactly. I’m not sure if there’s some setting I need to tweak, but if anyone knows, I’d really appreciate the help.
By the way, the workflow is excellent ;)
can someone update this? apparently some nodes have are no longer available for download.
Hi! Thank you for this amazing GGUF All-in-One workflow. It works great even with heavy models!
I have a question regarding the audio settings. In the standard LTX-2.3 workflow, there is an "audio_start" parameter to specify the starting point of the music file. However, I couldn't find a similar setting in this GGUF version.
Is there a way to specify the audio start time (offset) within this workflow? Or do we need to trim the audio file externally before loading it?
Thank you for your support!
I suggest you take a look at "10eros's" workflow; it incorporates nodes for secondary audio processing that can significantly enhance audio quality. Additionally, I hope you will explore LTX's seamless stitching capabilities—for instance, a feature that takes the final second of one video clip and the first second of the next, generating transitional frames in between to create a smooth splice. Of course, what we ultimately need is not just the ability to connect two clips, but rather the capacity to seamlessly stitch together multiple video segments simultaneously. If you are familiar with "WAN VACE," you know it offers precisely this functionality—though it currently lacks audio generation support. You can search for this specific workflow project right here on the site; it is designed to handle the batch stitching of multiple video clips, and I encourage you to download and study it. I am confident that LTX can bring this ambitious project to life!
@Silicon_Mirage Thank you for your suggestions. Unfortunately, I hardly have any time left to work on further updates. Sorry.
i get a error on this LTXVConcatAVLatent saying TypeError: 'NoneType' object is not iterable. not sure what this is.
also where the workflow tells you the spot to change video length, the bow is completely blank and there is nothing there to edit
"It looks like a version mismatch. Try updating your custom nodes via ComfyUI Manager and restarting. If the fields are still blank, you might need to delete the node and re-add it manually."
@sanmarinogmip394 i tried updating all the nodes and nothing worked. do you know which specific node is that one in?
Thanks for the reply!
Regarding the blank fields, the node responsible for video length is usually the "Video Generation (LTX-2.3)" node or the "LTXV_Config" node. If these fields are invisible, it often means the "LTX-Video-Helper-Suite" or "ComfyUI-KJNodes" isn't properly loaded. I recommend double-checking those specific custom node packs in your Manager.
As for the TypeError: 'NoneType', this happens when the LTXVConcatAVLatent node receives an "empty" input. Please ensure that both the Video Latent and Audio inputs from the previous nodes are actually generating data and are correctly connected. If the generation fails halfway, this error will occur.
I hope this helps you find the right node!
If these steps still don't resolve the errors, there's a possibility that your workflow file itself has become corrupted during the updates.
In that case, I recommend discarding your current workflow and re-importing a fresh copy of the original JSON file. Once re-imported, please follow my setup guide starting from the 1:41 mark to ensure the initial configuration is correct, and you can also refer to the 3:20 mark for switching modes and resolution settings.
Starting with a "clean slate" is often the fastest way to fix these persistent issues. Good luck!
Could anyone help me with using a reference image to go with my prompt? Super Newby I know but I have the workflow running fine but I can't figure out what to turn on in order for it to register the reference image to work with the prompt.
I noticed some people are having trouble with the Image-to-Video/Reference function. I've covered the detailed configuration in my walkthrough video.
The explanation for the reference image part starts at 3:20. Feel free to take a look if you're stuck!
3:4 accept ratio don't work