With the native comfy implementation of Hunyuan I have tweaked the workflow to work for 12GB VRAM cards. It looks like you can get at least to 73 frames and probably a bit more. It takes about 8 minutes for a 4070Ti to run 20 steps.
Do make sure to update your comfy and to get the exact result above the guidance I put up to 10 (I reduced it for the base workflow as it caused some burning for some prompts).
As its not as complete as the wrapper node there is a few less features than that for now. But there is also no crazy special installations that you need to do.
Links to model downloads here:
https://comfyanonymous.github.io/ComfyUI_examples/hunyuan_video/
Description
FAQ
Comments (170)
BEST !
Perfect timing for Project Odyssey! I was just about to dive into Hunyan. Also Congrats on 12GB reduced from 16GB from Kijai
I keep getting "EmptyHunyuanLatentVideo" is missing
You need to update Comfyui and all the custom nodes involved
@epsilon9 Ended up being related to the desktop version of ComfyUI. Using portable worked.
@epsilon9 eh question: how do I actually update comfyui (desktop app)? update all nodes in manager does nothing for me
@marcouscousaurelius It will be pushed there eventually - apparently its not on the desktop standalone yet.
install this node, it has the node you need in it -- ComfyUI-HunyuanVideoWrapper
still missing , though i installed that HunyuanVideowrapper ???
Using the portable standalone with Kijai's nodes installed, I also don't have this node, and Manager can't locate what's missing. Using the webui.
Fixed it. Manually download https://github.com/comfyanonymous/ComfyUI/blob/e4e1bff60532ea1a2e2550a1d9beb9b87bfd8c7c/nodes.py and put it in ComfyUI\custom_nodes\ComfyUI-HunyuanVideoWrapper and then https://github.com/comfyanonymous/ComfyUI/blob/e4e1bff60532ea1a2e2550a1d9beb9b87bfd8c7c/comfy_extras/nodes_hunyuan.py and put it in ComfyUI\comfy_extras.
@mwoody450 this fixed my issue too but now i'm getting a new problem, it says "ValueError: too many values to unpack (expected 4)" when it gets to the sampler node.
@mwoody450 thank you very much, this solved my problem too
Where is the download link for the new version of the model?
is this it?
tencent/HunyuanVideo at main
Is this for lower VRAM?
how can I add lora with this workflow
how to add lora to gguf?
can it fit in 10gb?
It looks like it can, start with a low frame count like 10
@Rating_Agent Yes, try lowering the resolution for faster generation. But 10gb works fine. Try 512x416 for example
but can it fit 6gb lol
nice 1. Hunyuan has so much potential, thanks for sharing.
how much system ram on top of the 12GB vram
32GB
Is it correct, that this workflow already converts FP16 to FP8?
Also, how would one incorporate the Huanyun Fast Video, that released recently?
Another question: Can one slow down the animation or is it all depending on the prompt?
Thanks for this great workflow.
This is the best workflow I've found, even better than the low VRAM flow from Kijai. I'm running it with the fast FP8 model and getting great results. The Kijai flow was somehow returning very blurry outputs. I run it with a Guidance of 10 and a Shift of 17 and I've gotten really great results. Thank you so much for sharing this, really appreciate it!
How do you set it up to run the new Fast model?
@lost_moon I just swapped it for it tbh, nothing more. I use the fast model from Kijai: Kijai/HunyuanVideo_comfy at main
The one that says "fastvideo".
I think the only other modification I made is to use the Clip-Vit-Large-patch14 instead of Clip-L, it makes gens faster when changing the prompt. I posted a couple of (gay) pictures to the gallery. Maybe you can get the workflow from them?
@DomDomTomTom Hunyuan Fast is freaking amazing. Wow
Thanks for your insights. Do you know what Shift does? Or where the docs are that describe it?
@SteveWarner My understanding of Shift (and someone feel free to correct me if I am wrong) is that it allows the model some flexibility from the original output. So it helps add more detail and variation to the final output by letting the model explore a little. Kind of an equivalent to CFG scale in a way I guess?
hunyuanVideoSafetensors_fp8VideoModel.safetensors I have this file and had to put it in unet folder to find it, I have Clip and VAE as per the page and I get 1 iteration in 70 seconds on an RTX3090 I tried just 1 frame and it still said step 1 of 20 and was so slow, what am I doing wrong?
@rogue8888992 Yeah that's not right. Try and share your workflow so maybe someone can help. It's difficult otherwise. Also make sure your Comfy is up to date.
@DomDomTomTom I've got hunyuan_video_t2v_720p_bf16.safetensors now so I can literally drop "Hunyan 12GB Workflow.json" on and press Run, I've updated ComfyUI yesterday and also all plugins, it appears to work, if I drop resolution to 208x208 and 5 frames and 20 steps it works at like 1s/it - I do notice even running this 12GB workflow my GPU ram usage goes from less than 0.5GB to maxed out at 24GB - is graphics drivers a possibility? they're maybe a couple of months old
at default settings for the 12GB workflow First iteration completed at 264s/it - I checked GPU drivers, they're from August, is that likely to be an issue? SD, SDXL & Pony run fine, 25 steps at say 3it/s or faster depending on resolution
I keep getting "EmptyHunyuanLatentVideo is missing" even though I've updated comfyui and the custom Hunyuan nodes. What am I doing wrong?
update comfyui
@Rindon like i said in the post, comfyui has been updated. the error still persists.
@catsinhats58 Try again? I had the same error, ran Update All, and no longer have the issue
I had this and fixed it by running update_comfyui.bat
would it brak if I update it to do image to video?
Does it accept LORAs?
It should yes - there are several already. They need more than a 4090 to train though.
@Inner_Reflections_AI how do we add lora to this workflow
@gigsa search for the lora loader node and just hook it up
In the Node "DualClipLoader"...under "type", it can't find "hunyuan_video". The dropdown only lists; sdxl, sd3, flux.
How do I get "hunyuan_video" to be recognized by "type" within that node?
in the manager update the custom gguf node package
@boyetosekuji how? What is the thing I'm supposed to update called? I'm confused and I've updated everything and still having the same issue as OP.
@catsinhats58 ComfyUI-GGUF
What's different about this workflow compared to the official Comfy one? Is it just the final ending node that makes it compatible?
depends what you mean by "official" the "official" hunyan setup requires something around 60-70 Gigabytes of VRam. A commercial 4090 has 24 and even an industrial grade ADA RTX 6k "only" has 48 so you would need TWO. Then there is Kijai's destill on github which gives you about 2 minutes on a 4090 at 30 frames. Then there is this. Its MUCH more memory efficient and for some reason doesnt even have forced unload which saves a lot of time you otherwise spend loading and unloading models whilst generating frames. The funny thing is it basically uses the same models. There is the FP8 version of the video model, the clip-vit-large clip model, the VAE BF16 model... well it uses llava_llama3_FP8_scaled textencoder model which... compared to the llava-llama-3-8b-text-encoder-tokenizer saves up major VRam resources, i guess. I guess you could use the text encode in the official workflow as well. Usually when it comes to LLMs i prefer 13b because 8b are not good conversation, then again its for painting and not chat. Yeah... the text encoder "large" language model basically. Very efficient as all models have to be loaded at once and a LLM is going to make a dent on your VRam if you are operating on commercial hardware. You could use forced unload on the models i guess to trade speed for runtime. Heh, its almost as we arrived in the 90s again trying to squeeze software into ram. Can't wait for the first fastloaders.
Sorry, what I meant was: what is different about the workflow in terms of the parameters in the workflow? When I referred to the "official" workflow, what I meant was the "official" workflow provided by ComfyUI, which looks the same to me as this one. I got it from the ComfyUI Discord in their official announcement about native hunyuan support. I was wondering if certain parameters in your workflow were changed to make it use less vram than the one ComfyUI provides, as I wanted to understand the settings better and to be able to understand more about what settings increase vram usage and which ones don't.
I'm not sure if you uploaded different versions or not, so to clarify, the one I'm talking about is in "hunyuanvideo12GBVRAM_v10.zip", and the workflow inside is titled: "Hunyan 12GB Workflow", apologies for the confusion!
Very strange issue with this workflow.
If I import the workflow and run it with all the defaults, it runs fine. But if I change the prompt to something else, I get an Invalid Buffer: 81GB error
How does workflow put the hunyuan lora?
thanks so much for sharing this!
Anyone tried this using an AMD GPU? I often have issues with different parts of ComfyUI and A1111 due to poor support for AMD.
I really wish the larger model was less than 25gb. I'm guessing it won't be usable even on 4090 or 7900
This workflow works on 12 GB graphics cards.
Hello, I got this error when generating a video finish while using your workflow; "replication_pad3d_cuda" not implemented for 'BFloat16',
Any way to make it work for img2vid?
This needs answering
This model is not Img2vid but the people behind it plan to release one in Jan.
@Inner_Reflections_AI Any update on this since the model has been released by now?
@HGBDN1510 Wan has gotten a lot more support and likely will be king going forward - hunyuan is still very good though
Was working, now getting "EmptyHunyuanLatentVideo is missing". I've updated all in Comfy. Any ideas?
You may have to update the comfyui base repo. That node is part of the base repo.
This workflow has worked far better than what anyone else has recommended, but is there a way to implement the new video enhance tools with it? It exports feta_args, which only connect to the HunyuanVideo Sampler nodes.
Great work! Thank you!
(4070Ti)
Even on an RTX 3060TI 8GB it works, great job. Thanks!
It takes 10 minutes to realize that the process fails in the VAE Decode because it run out of memory... I have the same card and I can only generate very low resolution videos, what size have you set?
I have 64 GB, which is probably why it's fine.
how long is it taking for you?
@KotatsuAi Try changing the tile size. You won’t have to generate it again it should just kick off from the decoding.
creativity of this model? you tried anything out of the ordinary?
Thanks for sharing, it’s the most stable workflow I have used.
damn i can use the full model with this. Any plans to add Lora support? is there a way to have negative prompts too?
did you try to add just the LoadLora node inbetween the LoadDiffusionModel and the model receiving nodes (Sampler, Clip)
i don't see setting with this workflow to use sageattention which was a massive improvement to my generation speed in a different workflow (spent ages getting it installed without error). is there a way to add it to this one or is it not needed with these nodes? also a 2nd issue, unlike otherworkflows this one doesn't have any lora input connections, how do you put a lora in this workflow?
nice!
It does not work on my 3060 12GB. It seems to make OOM for VAE, even with the lowest overlap & tile size (64) for it. Has anyone actually managed this with a 12GB card?
I have a 12 GB card and it works.
@Inner_Reflections_AI What sort of generation time or s/it for a 480x480 vid with 73frames and 20 steps?
@rogue8888992 5 mins roughly I think for that res.
I can confirm it works w/ my 4070ti 12gb
@Inner_Reflections_AI my 3090 bogged down and went way longer than 5 mins
how to fix?
Prompt outputs failed validation DualCLIPLoader: - Value not in list: clip_name1: 'clip_l.safetensors' not in ['clip-vit-large-patch14\\model.safetensors', 'llava_llama3_fp16.safetensors']
Firstly congratulations on the workflow!!! I'm having a lot of fun with it. I would like some help to know where to try to create slightly longer videos.. in the default configuration it is only being generated at 2 seconds. Could you give me a tip?
Increase the latent frames if you have enough vram - you can decrease resolution if you need to. Max is 200 frames with this model
I'm using a GPU server with an Nvidia a6000. When running the workflow, I'm running into an error with VHS_Video Combine "The minimum required Nvidia driver for nvenc is (unknown) or newer". Is it possible that VHS_Video Combine doesn't support the newer driver?
When loading multiple loras, from this method, https://civitai.com/articles/9584
It doesn't seem to work with your workflow because it wants me to use the HunyuanVideo Model Loader node over the Load diffusion model node. I can kinda get it working with normal lora though. Just curious if there's a better way of handling multiple loras on this workflow. Ty f
Hello, I have also been trying to get multiple loras working. If you dont mind could you share your workflow or just a screenshot?
@championgreat1239 Yes I found one, I just recently start using this one https://civitai.com/images/47386469 (just download the video and drop it in your comfy for the workflow but u prolly already know that.) I'd adjust the steps to 20. Idk how ppl get away with 8 steps but you can try it. When using two loras, it seems 0.6 strength on both is the sweet spot.
This is awesome, i can't believe i can run on my 8G 1070ti for a really nice quality, but it costs me 2.5 hours to make a 3 SEC video~~~
Thank you for this workflow! I was able to run this as-is—no changes to any parameters—with a 3060ti (8GB VRAM) and 32GB RAM. It took 19 minutes to render, but still!
Thats great, im using rocm on a 6700XT (12GB VRAM) and 64GB RAM and it took more than a hour to render, guess i really need another gpu for AI
Anyone able to get this working using AMD?
It works on Linux (Ubuntu).
Not on Windows. It crashes the driver. I run a Radeon 7800 XT with the latest drivers from december 2024.
It's taking me about 3 hours to do a 3 second video on a 4070ti - what the hell am I doing wrong? I must be missing some optimisations or something
Nevermind. A couple of stupid arguments I had in the command line were causing it. --Disable cuda Malloc and --Disable Smart Memory specifically.
@BellaMartinez Where did you include these command lines?
@AshleyArnes They were in the batch file
It would be awesome to see you make a new workflow with lora support!
Does anyone have loras working on a 3090? As soon as I add in a lora node to this workflow (or others...), my GPU memory usage shoots up to 27GB+ and ComfyUI just stalls out.
Yes works fine on my 3080ti, definitely something else wrong, use this method https://blog.comfy.org/p/running-hunyuan-with-8gb-vram-and
@GRIJAY thanks - turns out I needed to mess around with the comfyui commandline args a bit :)
@ultraduster45 What did you include in your command args and where?
same issue here
@AshleyArnes and @friendzcornerz810 if you're still struggling with this (sorry, didn't notice the reply earlier), you can try adding --lowvram or --disable-smart-memory (or both) to run.bat for the windows standalone. Alternatively, try a fresh install of comfy and re-setup workflows. Either of those worked for me.
It takes an hour on my 4080 with the default settings and even longer if i set it to fp8_e4m3fn and i dont know why
nvm i did a clean reinstall and now the test workflow runs in about 6 minutes
@AI_After_Dark Did you reinstall comfy ui as a whole? I can't even make a video in 10 hours with my 4080 and I tried like 5 workflows already, no success
@AshleyArnes I am using Stabillity Matrix so reinstalling CofmyUI is just a few clicks. Also the performance is not stable. Some videos still take 20 minutes and i don't know why since the prompt is almost the same length.
@AshleyArnes it is likely that you are using CPU if it is taking that long... try seeing how much VRAM is being used... or try HunyuanFast
@az420 I always click on the run_nvidia_gpu, when I start ComfyUi and during generations its on a 100% workload! Gonna check HunyuanFast too thank you
If you have correct torch i suppose you should google the nvidia drivers vram fallback comfyui. Basically your pc uses ram instead of VRAM
I tried this (and other) workflows too on my 4080 and it would take ages to generate a video. What am I doing wrong? I just installed ComfyUi, downloaded everything this workflow needed and it just does not generate anything, not even in 10 hours...
what does the console output say? what resolution and video length?
make sure you got the correct weights selected with the diffusion model, that did it for me
you used too many prompts, I tried many times, the more prompts you have the more Vram it takes, I'm sure you either copy those nice post or put yourself, try just one or two lines
My 4090 is working on all 24Gb of VRAM and it takes one hour to make something, while the pc barely could do anything else. What am I doing wrong? Than, why any prompt I try to make realistic, it generates just in anime style?
On my 4090, unless I'm using the Fast video checkpoint, I have to set weight_dtype to fp8_e4m3fn. Try that and you should see improvement. Without that the regular Hunyuan model seems to big for even a 4090.
@Synnyr I tried, but with fp8, the videos were really bad quality. I abandoned Comfy. On Forge I generate 3000 pixel images in less then 15 seconds. Too bad I can't create videos, but I can wait if maybe one day Forge will be updated.
If you have correct torch i suppose you should google the nvidia drivers vram fallback comfyui. Basically your pc uses ram instead of VRAM
Is there hope for my 2060 6GB?
No
No matter what I try I can't get comfyui to recognize my text_encoders
I ran it successfully
I ended up moving the text encoders to a different folder and it worked
Just generates pixelated noise, like TV static. Can't figure it out. Comfy is so complicated. 4th time attempting to learn this UI and there's just so many moving parts and one tiny little piece missing is all it takes. Impossible to tell what's gone wrong in order to fix.
This should be plug and play.
I also just get static no matter what I do, it runs though, just produces a static-filled bunch of moving colored pixels.
@Malikona Honestly I would start by just reinstalling comfy - this is as basic a workflow as possible if it doesn't work its unlikely to be the workflow itself.
@Inner_Reflections_AI Are there any other ways to do this? I'm getting interference as well. I just installed ComfyUI (from their repository) use these versions:
hunyuan_video_720_cfgdistill_fp8_e4m3fn.safetensors
hunyuan_video_vae_bf16.safetensors
4060 ti 16 gb
Me too :( I did not know how to search for this "fried" / "static" video problem. I am new to this kind of stuff, I tried to do everything properly but I am getting this distorted static video every time.
So yeah, I fixed it, you have to use the FP32 VAE, and the FP8fastvideo model, and clip_l for clip1. Any other configuration gives you the static.
@Minase460 Thanks for sharing!
try to delete ALL the models that you have installed completely. and download it again, it helped me. I deleted all the checkpoints, all the vae.
I've tried so many workflows including this one, and it takes 4 hours and produces just noise. I have a RTX 3060 with 12gb Vram and 64gb system RAM so should be well within spec. Anyone else have this?
Something is wrong with your system, with the 3060 12g this workflow works perfect for me. A turk put in legnt 1, the result will be an image but testing the whole process.
@papuchi79 I reinstalled comfy and it now works, not perfectly but it does work.
I am having trouble with the VAE loader. The start of the error message is "HyVideoVAELoader
Error(s) in loading state_dict for AutoencoderKLCausal3D:". It seems like the vae listed here is incompatible with the hunyanvaeloader, even though I downloaded and used the one mentioned in this tutorial. Any ideas why it can't load?
Very nice, your workflow in my 4070TIS, it only takes 4 minutes to render 120 seconds of video and the quality is very good, time is very much appreciated!
120 seconds WTF???? OMG?!!!
MystAI i think he meant 120 frames. 120 sec is nuts
I downloaded the workflow.
I followed the instructions where to get the models. (seemingly pointing to 16 bit checkpoints)
OOM on cuda 0
I found instead an 8 bit model to correspond to comfyui --fp8_e4m3fn-unet
OOM on cuda 0
I replaced the nodes that load VAE and text models with the nodes from multiGPU to place the text and vae models on cuda 1,2,3 devices.
Still OOM on cuda 0
I have 4 GPUs (old, but they each have 12.2GiB VRAM)
Supposedly the workflow should work.
Why isn't it?
Any suggestions are welcome, which nodes, which specific checkpoints I should use. Perhaps the urls here are outdated?
Thanks.
Are there text encoding models that can handle longer prompts and work with Hunyuan Video model?
"Token indices sequence length is longer than the specified maximum sequence length for this model (192 > 77). Running this sequence through the model will result in indexing errors."
Another post mentioned Long-Clip but I'm still getting the same error message: https://huggingface.co/zer0int/LongCLIP-SAE-ViT-L-14
Thanks for this workflow, Ran very well on my RTX 3060, Generation times about 6 to 8 minutes with 2 Loras loaded, great work.
@prodwaj Easy lora stack (comfyui-easy-use) > Apply lora stack (comfyui-jackeupgrade) then i just connected the needed link to model sampling, scheduler and clip nodes
What resolution and the amount of frames?
@Anselmo 720x448 and upscale to x2 that of the resolution 1440x896, 24 frames interpolated to 2.5 to reach 60 fps, 10 to 20 steps depending on what i'm generating
can you show me the resul ?
@xdon101993209 On my profile
First time I finally manage to run a workflow O.o I've been trying quite a lot of i2v Hunyuan workflows todays and can't seem to make any of them work. If you could manage to put together an image to video workflow as efficient as this one, that'd be lovely!
Cheers. Thanks for sharing.
I've been looking into this too. Based on what I've found, it's not possible with Hunyuan due to how it generates video, but I find it hard to believe there isn't some sort of solution. Custom LORA's trained on offshore servers seem to be the way to go.
You are using the base bf16 to do this? the 24gb model?
how using lora with this ?
This workflow instantly filled my 24gb of VRAM and started swapping into system RAM, becoming unbearably slow.
I noticed that my base model is called hunyuan_video_720_cfgdistill_bf16 instead of hunyuan_video_t2v_720_bf16. Found download, both files seem to be the same, at least size-wise. Any idea why its still chugging so much memory?
Dude this runs amazing on my 4070 12gb and 32 gb system ram. Only took about 10 minutes and great quality.
Works very well on 5070ti, Only mod I made is added lora to workflow. Total time is 8.5 minutes with default settings on workflow.
Not working. On launch i am getting all 32Gb ram oaded and then launch interrupted.
It takes my RTX 3060 12gigs about 30minutes to generate 20 frames with this. Not sure how I can get it down lower.
There is much more you can do with WAN these days. I would look for one of the new workflows there.
LTXV