🚀 MrXin LTX 2.3 I2V EROS V6
This workflow has continuously evolved based on user feedback to deliver one of the strongest LTX 2.3 I2V experiences available. Thx everyone for the feedbacks.
V6 is a major leap forward, focusing on stronger motion preservation, better prompt adherence, and more user control while keeping excellent stability and VRAM efficiency.
🔥 What’s New in V6 (Compared to V5)
Added LTX2 NAG (Negative Attention Guidance) on both First Pass and Final Pass for significantly improved prompt following and motion quality.
Switched to VAELoader KJ nodes for Video VAE and Audio VAE (better compatibility and stability).
Improved Final Pass by repositioning the LTXVCropGuides node for cleaner latent processing.
Added flexible sigma system: Manual Sigmas (8 steps) or automatically generated sigmas via LTXVScheduler + Sigmas Sigmoid (12 steps).
Final Pass is Manual Sigmas only for stability.
Audio volume default set to 0 (neutral) with easy slider control.
Enhanced image conditioning using LTXVImgToVideoInplaceKJ in both passes.
Cleaner workflow layout, better group organization, and more helpful in-workflow tips.
✅ Core Features
Extremely strong, natural and consistent motion from First to Final Pass
Excellent anatomy, lighting, physics and NSFW detail
High-quality 24 FPS video with perfectly synced audio
Easy model switching between 10Eros V1 FP8 and Distilled 22B
Built-in Video Editor (RTX Super Resolution + nmkdSiaxCX upscaler + RIFE 48/60 FPS)
Optimized for 16GB VRAM setups and 32GB RAM
📥 Required Models + Direct Download Links
10Eros V1 FP8 (Main Checkpoint) → ltx2310eros_v1_FP8.safetensors
Distilled Model → ltx-2.3-22b-distilled_transformer_only_fp8_input_scaled_v3.safetensors
Text Encoder → gemma_3_12B_it_fp8_e4m3fn.safetensors
Text Projection → ltx-2.3_text_projection_bf16.safetensors
Video VAE → LTX23_video_vae_bf16.safetensors
Audio VAE → LTX23_audio_vae_bf16.safetensors
Preview VAE → taeltx2_3.safetensors
Spatial Upscaler → ltx-2.3-spatial-upscaler-x2-1.0.safetensors
Model Upscaler → nmkdSiaxCX_200k.safetensors
Distilled LoRAs (First & Second Pass)
📁 Folder Structure
ComfyUI/
├───📂 models/
│ ├───📂 diffusion_models/
│ │ └─── ltx2310eros_beta.safetensors
│ │ └─── ltx-2.3-22b-distilled_transformer_only_fp8_input_scaled_v3.safetensors
│ │
│ ├───📂 text_encoders/
│ │ └─── gemma_3_12B_it_fp8_e4m3fn.safetensors
│ │
│ ├───📂 clip/
│ │ └─── ltx-2.3_text_projection_bf16.safetensors
│ │
│ ├───📂 VAE/
│ │ └─── LTX23_video_vae_bf16.safetensors
│ │ └─── LTX23_audio_vae_bf16.safetensors
│ │ └─── taeltx2_3.safetensors
│ │
│ ├───📂 latent_upscale_models/
│ │ └─── ltx-2.3-spatial-upscaler-x2-1.1.safetensors
│ │
│ ├───📂 upscale_models/
│ │ └─── nmkdSiaxCX_200k.safetensors
│ │
│ └───📂 Lora/
│ └─── ltx-2.3-22b-distilled-lora-1.1_fro90_ceil72_condsafe.safetensors
│ └─── ltx-2.3-22b-distilled-lora-384-1.1.safetensors
Custom Nodes Overview – MrXin LTX 2.3 I2V EROS V6
ComfyUI-KJNodes
rgthree-comfy
ComfyUI-easy-use
ComfyUI-mxToolkit
ComfyUI-VideoHelperSuite
ComfyUI-LTXVideo
controlaltai-nodes
comfyui_nvidia_rtx_nodes
GACLove/ComfyUI-VFI (voor RIFE)
comfyui_memory_cleanup + comfyui-impact-pack
Pro Tip: Use the new 10Eros V1 FP8 model with its recommended trigger words for the best results.
✅ Quick OOM Fix Guide
If you’re getting Out of Memory errors or weird faces in the final pass, follow these steps in order:
1. Update or Fresh Install ComfyUI + Install Missing Custom Nodes
Download and run the Easy Installer: 👉 https://github.com/Tavris1/ComfyUI-Easy-Install
Or watch this short video: 👉 https://www.youtube.com/watch?v=CgLL5aoEX-s&t=1041s
After updating/installing:
Open ComfyUI → go to Manager → click Install Missing Custom Nodes
Restart ComfyUI completely.
2. Add the Low VRAM Flags
Go to your ComfyUI folder.
Find the file ComfyUI.bat.
Right-click → Edit (open with Notepad).
Find the line that has ComfyUI\main.py .
Add the flags at the end so it looks like this:
bat
ComfyUI\main.py --lowvram --reserve-vram 6 --preview-method none --disable-xformers --disable-smart-memorySave and launch ComfyUI using this .bat file.
3. Increase Windows Swap File (Virtual Memory) This helps when VRAM still spikes:
Right-click This PC → Properties → Advanced system settings.
Click Settings under Performance → Advanced tab → Change (under Virtual memory).
Uncheck “Automatically manage paging file size for all drives”.
Select your main drive (usually C:) → choose Custom size.
Set:
Initial size (MB): 49152
Maximum size (MB): 98304
Click Set → OK → restart your PC.
After doing all three steps, reload the workflow and test with Longer Side 1024 and Video Length 20 seconds first. This combination fixes OOM for almost everyone on 12GB cards.
Version History & Previous Updates
V5 — Major base model upgrade to 10Eros V1 FP8. Significantly better realism, anatomy, lighting and motion coherence. Updated model loading and separate LoRA handling for each pass.
V4 — Introduced separate Manual Sigmas for First and Final Pass, much stronger Final Pass motion, full TwoWaySwitch system for model selection, expanded LoRA stack, and refined Video Editor.
V3 — Solved the biggest V2 complaint by preserving much more motion in the Final Pass. Added Model Upscaler (nmkdSiaxCX), option to disable Final Pass, distilled model support, and custom resolution.
V2 — Added full built-in Video Editor (RTX Super Resolution + RIFE to 48 FPS), improved LoRA suite, better audio sync, and major stability enhancements.
V1 — The original production-ready workflow. Dual-pass system, strong audio integration, excellent low VRAM optimization (12GB cards), and live previews.
Disclaimer:
This workflow is provided for entertainment, artistic, and creative purposes only. It may not be used for any illegal, harmful, non-consensual, or malicious activities. Please use it responsibly and respect all applicable laws and ethical guidelines.
— MrXin (May 2026) 🔥
Description
🚀 MrXin LTX 2.3 I2V EROS V5
This is the latest evolution of the popular MrXin LTX 2.3 I2V EROS workflow.
V5 introduces a major upgrade to the base model: the new 10Eros V1 FP8 checkpoint. This updated model delivers noticeably better anatomy, lighting, motion coherence, and overall visual quality compared to previous versions.
🔥 What’s New in V5 (Compared to V4)
New 10Eros V1 FP8 base model — significantly improved realism, detail and prompt adherence (see full details here)
Updated model loading with DiffusionModelLoaderKJ for better compatibility and performance with the new FP8 checkpoint
Separate LoRA loading for First Pass and Final Pass (optimized distilled LoRAs for each stage)
Improved RTX Super Resolution (now default 2x upscale)
Refined node layout and connections for better stability
Default video length remains 20 seconds with strong motion preservation
FAQ
Comments (107)
How can I make it faster if I have more VRAM?
is it possible to run your workflow on amd gpu's? without nvidia nodes like rtx upscalers. I have RX 9060 XT 16gb
i have 7800 xt 16gb, deleted nvidia upscale, and it worked
@MokiMok31 For me, the generation refuses to go past the Negative prompt stage. I waited for about 20 minutes, but it didn't budge. It did proceed once, but at the First Sample, it simply errored out and then "Reconnecting". Everything is configured correctly in the workflow. I have 48gb pagefile and 32gb ram.
@MokiMok31 how did you make it run? any special params? because I'm using default only "--preview-method auto --enable-manager-legacy-ui" launch options. Or you mean you've deleted nodes in the workflow aswell? Because I removed the entire Video Editor nodes from the workflow, and also didn't download the rtx nodes. But that didn't help. I have Re-bar enabled in UEFI, and the GPU is not overclocked or undervolted. And my generation just pauses on (1) 37% - Negative here:
got prompt
VRAM清理完成 [卸载模型: True, 清空缓存: True]
RAM清理完成 [16.1% → 15.0%, 释放: 380MB]
VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
VAE load device: cuda:0, offload device: cpu, dtype: torch.float32
Requested to load VideoVAE
loaded completely; 13982.92 MB usable, 1384.94 MB loaded, full load: True
CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
Requested to load LTXAVTEModel_
loaded partially; 12876.12 MB usable, 12313.62 MB loaded, 13127.87 MB offloaded, 562.50 MB buffer reserved, lowvram patches: 0
@wemite Got the same issue here, on RX 9070 XT 16GB vram and 32GB ram. You found another solution?
@Bitzito nope, it seems that ltx 2.3 doesn't work on amd properly... I tried other models and workflows (full and gguf), nothing works. wasted money for amd( I hope someday the ROCm amd drivers will work with ltx
@Bitzito I have a Gigabyte Aorus Z490I Ultra motherboard that supports only PCI-e 4.0 and lower (while RX 9060 XT is PCI-e 5.0). So I've set PCI-e 4.0 instead of AUTO in my UEFI settings and it hepled a little for WAN 2.2 gguf, idk why, but it generates a still image, not a video. It's already success for me (never even worked before). Now I need to test it on LTX
@wemite Yes i also gave up, ill stick to images for now.
@Bitzito I swear, LTX works. Try 10eros fp8mixed_learned model, set weight_dtype to fp8 e4m3fn. Memory settings - auto (not highvram or lowvram). Disable pinned and smart memory. I also use gemma-3-12b-it-heretic-v2_nvfp4.safetensors as a text encoder, and I've set the pagefile to 96gb instead of 48gb. Also set the "Chunk FeedForward" to something like 12/512 instead of default 4/4096. I had to wait a bit for Negative and Positive, but the generation speed seems okay for amd. But it's still very slow compared to even 3060 ti
Dear Author, thank you for such an excellent workflow. I use it extensively. HOWEVER! It contains an error—an error that is also present in version 4. You are sourcing the Audio VAE from the ltx2310eros_v1 checkpoint. Yet, when switching to a configuration utilizing a different model, Node #190 (LTXV Audio VAE Loader) continues to pull the Audio VAE directly from the checkpoint.
Consequently, if a user has loaded a different model—for instance, DasiwaLTX23Lightspeed_treasurechestV1.safetensors—and has either deleted ltx2310eros_v1 or never downloaded it in the first place, they are unable to select an external VAE, such as LTX23_audio_vae_bf16. Your node currently sources the VAE exclusively from the checkpoint. If this behavior is indeed an oversight, you might consider utilizing an alternative node—for example, the "VAELoader KJ Audio" node.
ow okey, thx for the info. i didn't know this was a thing. I wil try to fix it in the next update.
Thank you for providing the workflow. After generating a few videos with version v5, I found that the image quality is lower than the original images. Changing the multiple value to 1.0 results in an "out of memory" error.
I seem to be getting an outOfMemory error even with the multiple value set to 0.5, even with 16GB VRAM
Good Job.
But there are small errors:
The nodes "Diffusion Model Loader KJ" and "LTXV Audio VAE Loader" search for the models in the wrong path.
Im getting this error. AssertionError: SM90 kernel is not available. Make sure you GPUs with compute capability 9.0. This is related to SageAttention, but i dont see the node in the workflow?
Disable sageattention in the "Diffusion Model Loader KJ" node, click the "sage_attention" line and click "disabled"
That was it! Thank you!
Great workflow! Can you update your T2V workflow for the latest 10Eros too?
EDIT: I don't mean to sound ungrateful. Your work is outstanding and I thank you.
Also, I'm not exactly sure what the cause is, but there is a face drift issue. When I create a video using a photo of an Asian woman, she gradually changes to look like a Westerner. This happens with all the videos. Is there a solution?
Do you need to specify races more in LTX? I've noticed the same thing, maybe we need to specify?
Using i2v, getting distorted and blurty teeth during dialouges. Any idea how to fix that?
The workflow is really good. I'm getting some really nice results.
But I have a question. is there any way to generate the videos without sound? It's pretty cool, but not every video needs audio. I tried bypassing , the nodes m disconectiong audio vae, but it doesn't work. It's probably something really silly though XD
You can use video freeware like Handbrake to re-encode the video without audio once its done. The other way you could do it is through prompting, maybe.
getting very low sound with v5. Any fix ?
There is node Audio Adjust Volume inside Processing Video subgraph.
MrXin, you're very talented, and thank you for the workflow. It's very neat.
I'd like to ask you, in a future version or maybe even a separate workflow, if you have the opportunity, to create a V2V feature to create voiceover for videos, say for example, from a WAN.
Great and fast workflow! Any plans to make a version optimized for 32gb vram? Im unsure what the ideal settings would be myself
This workflow has been great! Getting fantastic results. May I ask what peoples runtime has been. I have quite a powerful Vram and my generations are taking 15mins sometimes, is that normal. Just want to get a rough idea of what the majority are getting
Can we get a version of this workflow with custom audio input?
This was working great, but suddenly my resolution and length sliders aren't showing up so I can't change those. Any idea how to fix that? Thanks.
Im getting subtitles when i prompt the characters to say something. Putting "subtitles" and "text" in the negative prompt doesnt work sadly. Did anyone have the same issue?
I've noticed the same. It rarely happened with the last workflow and models, but now it happens in >80% of the videos.
thanks for the workflow <3 looks super clean, but I am not sure what I am doing wrong but no matter what I do but with the v5, the second pass always generated a brownish video with a slightly different character and all foggy, any help? ty
Amazing work but being an utter noob I can't work out where to place the missing Rife in order for it to stop kicking out errors (rife-flownet-4.13.2 right?) - Can anyone help please?
Does anyone know why my vram usage is spiking in the prompts node?
Great workflow, but I've noticed that enabling the chunking causes the video to become a completely blurry mess. I'm using eros with the correct distilled loras. Any ideas why the chunking causes this?
SAME. Anyone know why?
Same. While the overall texture remains largely intact, distortion and blurring occur. If you turn it off, these issues do not occur.
Im getting the Error
"Error
No setter found for Height(easy getNode)"
How can I fix this?
This workflow seems to run slower than a WAN2.2 I2V workflow. 380s vs 275s. The results are pretty similar, minus the audio. Is there a way to speed it up? I thought LTX was supposed to be faster than WAN.
Disable the final video node
@Yourmomd yeah, I toggled off a few of the features there were toggles for, and got generations down to around 215s. I feel like if I had more than 32GB of RAM, it would be even faster, as it seems like the generations are fast, but the in-between loading steps take a long time.
I'm getting some issues, My videos are MUCH better looking with the just the comfy template so I'm sure it's ME doing something wrong. Should it look similar to the comfy template with the 10eros checkpoint? And also, isn't that c checkpoint? Am I suppsed to use that as a diffusion model in this workflow?
There are too many issues. The path for “LTXV Audio VAE Loader” is incorrect. The current path is set to the “checkpoints” folder.
thanks for pointing that out, it was driving me nuts
Hi everyone, can anyone tell me how to make breasts grow, or where I can get some Lora on breast expansion? I try to make this only by prompt, but with no result
Help! Why am I getting garbled subtitles when I add Chinese dialogue prompts? Is there a solution?
Install the NAG node, then add text-based subtitle-related prompts to the negative message list. The subtitles will practically disappear; it's at least 99.99% effective. You can find tutorials online; many people have shared their experiences because this node is indeed very useful.
When I look at the example videos, I find your workflow good, when I try it, it never really is that good, I'm probably doing it wrong or missing something important, but anyway... I'm actually having better cleaner results with my personal crappy messy workflow that is far less polished than yours, but something is not clicking on my pc at least (even if your example videos are really nice)
Sorry, I'm a total beginner but this is working great for me so far. Is there any way we can bypass saving the video which doesn't contain audio? and is there an option to select 720p or 1080p?
can you stop pinning nodes plz, very frustrating. it maked hard o debug and edit
any chance we might be able to integrate the First and Last frame stuff, or would that have to be a different setup completely? I've tried manually merging the base first and last LTX2.3 workflow with this one, but it's not quite working.
what could be an issue? I can't choose loras and can't move half of nodes. can't replace them. some of them work weirdly, e.g. require audio vae but when I click it seeks for checkpoint.
why do you keep renaming nearly all, if not all the models you use?
V5 is the best yet. Absolutely no issues. Smooth a silk even on 12GB. Amazing work!
How hard will it be to add end frame support?
Anyone have any luck using or implementing ID Lora for voice control?
also interested in this. How does this work? I've never attempted ID lora for voice yet. have you used it in other workflows?
I can see by the 0 responses over the last several days to everyone you probbly won't respon but I'll try I guess.
The final pass changes the animation of the first pass DRASTICALLY, the first pass has better animations about 90% after about 30 renders.
The final pass DEFINETELT makes everything look better for sure, but loses the better animations of the first pss.
I don't see a way of denoising the second pass, could you please help? if it's not too much trouble?
It's a fairly simple question, however i may just be ising something
The workflow is excellent. I've tried many, and this one is the most effective, both in terms of image and audio quality.
I'd like to share my Lora combinations. Additionally, I'll summarize some minor issues I encountered. I hope this workflow can be improved further, and I thank the author for their dedication and effort.
(1): Lora Combinations (Lora Name + Weight Intensity)
DR34ML4Y 0.5, furry 0.2, Licon-VBVR-I2V-390K-R32 0.8 (models "10Eros_v1").
I conducted extensive testing on some individual Loras, comprehensively testing them in terms of motion, image quality, and sound effects. I analyzed their advantages and disadvantages and found a combination that best suited my needs.
Because NSFW's Loras and models can cause some contamination due to various reasons, try not to stack too many Loras and high intensities, otherwise it may ruin the final result.
(2): Summary of Some Issues
A. The 1-pass result is completely normal, but more issues arise after the 2-pass. A grayish appearance occurs when the image background is dark; this issue is reproducible. The problem does not occur with a warm, white background.
B. For multi-person images, lip-syncing for speaking works perfectly in the 1-pass. However, after the 2-pass, sometimes multiple characters speak simultaneously, even though only one voice is heard. Even with additional descriptions such as gender and clothing, it's still impossible to accurately identify a specific character speaking.
Finally:
The workflow is excellent. If a NAG node and a Prompt_Relay node are added in future versions, this workflow will reach a new level. Thanks again to the author, and I look forward to the next update.
Thank you so much for the kind words and detailed feedback! 🙏
I'm really happy to hear that the workflow is one of the most effective ones you've tried — that means a lot.
Regarding your LoRA combination:
DR34ML4Y 0.5 + furry 0.2 + Licon-VBVR-I2V-390K-R32 0.8 on 10Eros_v1 sounds like a very solid combo. Thanks for sharing your testing results! I'll definitely try this stack myself. You're right about not stacking too many NSFW LoRAs at high strength — it can quickly cause contamination and degrade the final output.
About the issues you mentioned:
A. Grayish tint on dark backgrounds after 2nd pass This is a known behavior with the current Final Pass settings. The 2nd pass tends to slightly desaturate or cool down dark scenes. Possible fixes I'm considering:
Adjusting the Final Pass strength (currently 0.8)
Tweaking the sigmas or adding a small color correction node after the Final Pass
B. Multi-person lip-sync / speaking confusion in 2nd pass This is one of the harder limitations of LTX 2.3 right now. The model sometimes struggles to keep individual character audio attribution after the refinement pass. Adding more specific descriptions (gender, clothing, position) helps in the 1st pass but gets diluted in the 2nd. I'll look into ways to improve this.
Future improvements Great suggestion about adding a NAG node and Prompt Relay! I agree these would make the workflow even stronger. I'll try to implement them in the next version.
Once again, thank you for taking the time to test thoroughly and share your findings. This kind of constructive feedback is incredibly valuable.
Appreciate your support! 🔥
— MrXin
@MrXin Thank you for your recognition. Your work is truly outstanding and must have taken a lot of effort. I believe the subsequent versions will be even more perfect. Thank you again for your hard work.
Can't seem to get this to work.
On a 5080 + 32GB system memory and it constantly gives "out of memory" errors
The VRAM cleanup node says the required input is missing but it seems to be installed.
Hoping someone else knows what to do here
I never get that error and I do 30 seconds 1080x1920 on 5060 and 64gb system memory
This model uses a lot of Ram and Vram so whats happening to you is that it's getting your Vram to full capacity and it's using your Ram as shared Vram and because you only have 32gb of Ram you are running out of memory.
I can run this model with no issue with 8gb of Vram and 64gb of Ram. I think you need to get more Ram to run this model.
@piconejo Yeah that seems to be the case, was hopeful since the examples state it can run on my exact specs but I've had to run it with launch args to get a result.
If anyone else has this problem these launch args should at least get it to run " --cuda-malloc --disable-pinned-memory --normalvram --reserve-vram 8"
i ah yes, i have mention this in the model info near the end.
im getting low volume audio, can some1 tell me why is this happening?
Does this workflow need more time? Seems like having two runs with 2 distilled lora increase my generation time from 2 minutes (regular ltx2.3) to 5 minutes (this workflow). Is the extra time worth it or am I just doing something wrong with the workflow?
It outputs - Video/Image and Audio+Video. How do I output Video only and audio seperately?
Also quiet audio here sometimes
How to increase STEPS from 8??!! I can't find the node
You can consider the number of steps as the number of digits entered in Sigma's xxx Pass minus 1. This workflow consists of 8 steps in the First Pass and 3 steps in the Final Pass.
There’s an issue with the distilled LORAs. Unless I reduce the First Pass strength to 0.5, videos are a blurry mess with artifacts everywhere. But with 0.5 the quality is then really bad. Don’t know how you guys manage to have normal results. Maybe you don’t use photorealistic pictures?
any luck? My renders are a complete disaster too. blurry, horrific audio and then some.
EDIT: I fixed my issue by adding an additional distilled lora on powerloader.
Update is coming to hopefully fix this for you guys. I'm trying alot of stuff but it takes time.
This is so well designed as a workflow. Brilliant instructions and the best touch of all - which I think is inspired - is to have a section of all well used Loras and their respective TRIGGERS - BIG BIG THANKS!!
ModuleNotFoundError: No module named 'sageattention'
What is this and where do I find it?
bro, right there there in the Diffusion Model Loader. first node
@WhiteRider yes I know where it is in the workflow but where do I find it to install. I got an error saying i don't have it.
@alexandriaadrienne940737 install comfyui with the links in the worflow and follow the video on youtube.
@alexandriaadrienne940737 ah gotcha.. if you are on linux you can "pip install sageattention" its pretty simple, check youtube vids.
Outputs are blurry and completely useless. anyone get a fully clean render yet?
EDIT: Okay so I got a decent result but I had to add an additional distilled lora to the Power Lora Loader. Try adding the ltx-2.3-22b-distilled-lora-384-1.1.safetensors or something similar and see if that helps you all.
yeah thats what i recommend in the workflow on the left side, folder structure.
Thank you. I have been trying to figure this out all day.
Hi, I'm using 10ero FP8 with an RTX 5080 and I'm getting a memory error.
You still need more than 32gb of ram to run this model
Appreciate the workflow. What is the video editor portion of the workflow for?
Its for upscaling the resolution and the scale up the frame rate to the frame rate you want.
cant get it to work in anyway gemini says it a problem with SageAttention and windows
use the link in the workflow to install easy comfyui. There is also a video explaining how to install SageAttention
@MrXin you dont know how much this helped me thx
@MrXin can u make a txt to vid version and a video to video i woude try doing it alone but there is no way in hell that i know how to do it without going crazy and schizo
@sadsshit I will make these but will it take some time.
@MrXin and i will start loving you harder
What is this absolute dumpster fire of a workflow? You have no idea how these models work. You're loading the wrong Eros FP8 (the garbage unofficial one), the regular LTX distilled, TWO conflicting LoRAs when the creator of Eros specifically says to avoid 384, and fucking DR34ML4Y?
And to top it all off, you're using some obscure ass unnecessary nodes that are solved better by popular, common ones.
WARNING: This is a voodoo mix of all the wrong shit for morons, by a moron. Avoid this at all costs.
You are absolutely right, I am a newcomer to creating workflows. This got out of hand when someone once asked which workflow I used to make my ltxv videos. So I posted this one, never thinking it would catch on so well. I am always open to change, but I don't always have the time to figure everything out properly.
@MrXin Don't pay attention to the negative comments. Most of us are appreciative of your work. We're all learning together. 👍
touch some grass bro
V6 is using the right fp8 and best version of it. The civit eros fp8 got updated to fp8_mixed on like day 2. Its the old transformer-only one thats not ideal because it's more scuffed fp8. 384 is only more intrusive and base model aligned; it can change output and give plastic looks and is also a huge file size, gives choppy motion blur, but it's run at reduced strength I assume so it's more normalized. 72_condsafe is made to just run higher strength without the distilled distortion and to give a bit more motion detail and it also gives a different output look in general without burned skin details and stuff, but recently I find that you really want a full 4 way distilled lora block split to both first pass audio/video and upscale pass, keeping audio at 1.0 or 0.9 and video at 0.8/0.6 ish on the passes. All you gotta do is just load it into workflows, they have plenty of customization thats the point. Dr34ML4Y just seems redundant as a tack on lora but in some prompts seeds it could give an alternate motion but its a generalized lora used on an already generalized model I prefer OMNINFT 2.3 conversion if there is one lora you would run on every Eros prompt/image since it doesn't interrupt nsfw motions or sound, and it smooths finer details or can animate background effects or add different effects itself.
Is there a way to add a pre made Audio too instead of only an Image to generate a video? Or only Image is possible?.
This workflow is great at fitting within the 12 GB VRAM envelope but I felt there was a huge dropoff in terms of visual accuracy, graphical fidelity, and form compared with the official workflow. But then again the official workflow crashes a lot on a 3080 Ti with 12 GB VRAM, 64 GB DDR4 RAM, and 32 GB Linux swapfile.
Some extra work I had to do was install sageattention, create a symlink for the audio VAE from the vae to the checkpoints folder, and lastly replace the default first and second pass loras with the author's recommendations.
This workflow is amazing! Could you add voice cloning or premade audio in the next update?
Having trouble getting a clean video output. downloaded all the required models indicated on the left pane and video output is absolutely a distorted picasso mess. Any input would be appreciated
update: v6 automatically fixed the issues I was initially mentioning. huge props to mrxin for this workflow, absolutely crazy work