Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910879733252096 |
|---|---|
| author | Liu, Feng Zhang, Shiwei Wang, Xiaofeng Wei, Yujie Qiu, Haonan Zhao, Yuzhong Zhang, Yingya Ye, Qixiang Wan, Fang |
| author_facet | Liu, Feng Zhang, Shiwei Wang, Xiaofeng Wei, Yujie Qiu, Haonan Zhao, Yuzhong Zhang, Yingya Ye, Qixiang Wan, Fang |
| contents | As a fundamental backbone for video generation, diffusion models are challenged by low inference speed due to the sequential nature of denoising. Previous methods speed up the models by caching and reusing model outputs at uniformly selected timesteps. However, such a strategy neglects the fact that differences among model outputs are not uniform across timesteps, which hinders selecting the appropriate model outputs to cache, leading to a poor balance between inference efficiency and visual quality. In this study, we introduce Timestep Embedding Aware Cache (TeaCache), a training-free caching approach that estimates and leverages the fluctuating differences among model outputs across timesteps. Rather than directly using the time-consuming model outputs, TeaCache focuses on model inputs, which have a strong correlation with the modeloutputs while incurring negligible computational cost. TeaCache first modulates the noisy inputs using the timestep embeddings to ensure their differences better approximating those of model outputs. TeaCache then introduces a rescaling strategy to refine the estimated differences and utilizes them to indicate output caching. Experiments show that TeaCache achieves up to 4.41x acceleration over Open-Sora-Plan with negligible (-0.07% Vbench score) degradation of visual quality. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_19108 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model Liu, Feng Zhang, Shiwei Wang, Xiaofeng Wei, Yujie Qiu, Haonan Zhao, Yuzhong Zhang, Yingya Ye, Qixiang Wan, Fang Computer Vision and Pattern Recognition As a fundamental backbone for video generation, diffusion models are challenged by low inference speed due to the sequential nature of denoising. Previous methods speed up the models by caching and reusing model outputs at uniformly selected timesteps. However, such a strategy neglects the fact that differences among model outputs are not uniform across timesteps, which hinders selecting the appropriate model outputs to cache, leading to a poor balance between inference efficiency and visual quality. In this study, we introduce Timestep Embedding Aware Cache (TeaCache), a training-free caching approach that estimates and leverages the fluctuating differences among model outputs across timesteps. Rather than directly using the time-consuming model outputs, TeaCache focuses on model inputs, which have a strong correlation with the modeloutputs while incurring negligible computational cost. TeaCache first modulates the noisy inputs using the timestep embeddings to ensure their differences better approximating those of model outputs. TeaCache then introduces a rescaling strategy to refine the estimated differences and utilizes them to indicate output caching. Experiments show that TeaCache achieves up to 4.41x acceleration over Open-Sora-Plan with negligible (-0.07% Vbench score) degradation of visual quality. |
| title | Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2411.19108 |