This is a proof of concept workflow around simple temporal upscaling.
Every video model has some blurring when it comes to high motion scenes - both closed and open source. Depending on the model this can be at higher or lower motion with newer models being better than older as a general rule. Some styles tend to make their blurring effects at lower speed - especially anything with lineart, which could be due to training being more on realism or the way latents are compressed temporally.
One way people try to solve this is by upscaling and rediffusing over the video which definitely helps but in my experience at least there is a definite limit to how much it can help. This is where temporal upscaling comes in.
The idea around this is not a new concept per se for example MAINodes https://github.com/matlowai/ComfyUI-MAINodes attempts to sort out which frames need to be temporally upscaled and rediffuse over them. I took this to a logical conclusion and considered what if you just 'deroped' aka temporally upscaled the whole scene rather than trying to be picky.
I do think it has several advantages:
1/Your final result tends to have more motion consistency in the small details than with varying how you denoise
2/You can finetune the denoise to your result.
3/I am finding a reasonable result with simple 2x temporal upscale.
Of course the main disadvantage is that you are diffusing over a video that now is 2x the size and potentially upscaled at the same time. This is were speedup tools come in - low step lora as well as sol attention (which provides increasing benefits the longer the video is).
The purpose of this workflow was to create a workflow as close to base comfy. I use Kijai's node pack and VHS Suite to help manipulate the frames. The only real outlier node is the one used to slow down the audio for the temporal upscale. Feel free to encode empty audio if you prefer or find your own pack that has something which does this.
I simplified my workflow to remove anything not essential to demonstrate the concept. I expect you to add it to your own workflow or build upon this as a base. My tips when it comes to using this for a while:
1/If you are noticing some morphing especially in background stuff I suggest you disable sol attention at least for the base - it does do this to some shots but not others.
2/By default we create the base video at 0.5 MP at 20 steps - this is to get good sound and to speed to find a good seed, then we are temporally upscaling x2 and upscaling to 1 MP. Look at the yellow boxes to change the upscale settings.
3/Adjust denoise of the upscale to your video length and needs. Shorter and lower resolution videos often need less denoise whereas longer videos often need more. If you go too high your video will speed up and ironically become more blurry as a result. 0.45 is a good starting point but anywhere from 0.3 to 0.6 or higher is acceptable.
4/If you think the action is still to fast for a 2x upscale you can do 3x or 4x.
Description
Base Version
FAQ
Comments (5)
Thank you for this. I spent a little time with the MAINodes awhile back, and the technique seems promising, if compute-intensive.
I mean if you consider this a refiner/upscaler like it is my wf its not to bad as you dont have to temporal upscale everything. You can run 0.5 mp gens every minute and then take the 5-6 minutes to upscale if you like it for 10 seconds on a 5090.
I'm not sure I fully understand.
Does this just take longer than a standard upscaling, or does it also require more VRAM?
And if more VRAM, by what percentage?
Great work, by any means! ๐
You are doubling the length of the video and doing a low denoise pass to upscale it in slow motion. Because the motion is much slower the blurring goes away. It is much more intesive which is why we use speedup things to help and reduce as many steps as we can. The quality improvment is worth it though.
this metod is good for non 2d renders
2d art gets destroy in style
for 2d i recommend 50 steps with tea cache