Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Feng, Zhang, Shiwei, Wang, Xiaofeng, Wei, Yujie, Qiu, Haonan, Zhao, Yuzhong, Zhang, Yingya, Ye, Qixiang, Wan, Fang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910879733252096
author Liu, Feng
Zhang, Shiwei
Wang, Xiaofeng
Wei, Yujie
Qiu, Haonan
Zhao, Yuzhong
Zhang, Yingya
Ye, Qixiang
Wan, Fang
author_facet Liu, Feng
Zhang, Shiwei
Wang, Xiaofeng
Wei, Yujie
Qiu, Haonan
Zhao, Yuzhong
Zhang, Yingya
Ye, Qixiang
Wan, Fang
contents As a fundamental backbone for video generation, diffusion models are challenged by low inference speed due to the sequential nature of denoising. Previous methods speed up the models by caching and reusing model outputs at uniformly selected timesteps. However, such a strategy neglects the fact that differences among model outputs are not uniform across timesteps, which hinders selecting the appropriate model outputs to cache, leading to a poor balance between inference efficiency and visual quality. In this study, we introduce Timestep Embedding Aware Cache (TeaCache), a training-free caching approach that estimates and leverages the fluctuating differences among model outputs across timesteps. Rather than directly using the time-consuming model outputs, TeaCache focuses on model inputs, which have a strong correlation with the modeloutputs while incurring negligible computational cost. TeaCache first modulates the noisy inputs using the timestep embeddings to ensure their differences better approximating those of model outputs. TeaCache then introduces a rescaling strategy to refine the estimated differences and utilizes them to indicate output caching. Experiments show that TeaCache achieves up to 4.41x acceleration over Open-Sora-Plan with negligible (-0.07% Vbench score) degradation of visual quality.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19108
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
Liu, Feng
Zhang, Shiwei
Wang, Xiaofeng
Wei, Yujie
Qiu, Haonan
Zhao, Yuzhong
Zhang, Yingya
Ye, Qixiang
Wan, Fang
Computer Vision and Pattern Recognition
As a fundamental backbone for video generation, diffusion models are challenged by low inference speed due to the sequential nature of denoising. Previous methods speed up the models by caching and reusing model outputs at uniformly selected timesteps. However, such a strategy neglects the fact that differences among model outputs are not uniform across timesteps, which hinders selecting the appropriate model outputs to cache, leading to a poor balance between inference efficiency and visual quality. In this study, we introduce Timestep Embedding Aware Cache (TeaCache), a training-free caching approach that estimates and leverages the fluctuating differences among model outputs across timesteps. Rather than directly using the time-consuming model outputs, TeaCache focuses on model inputs, which have a strong correlation with the modeloutputs while incurring negligible computational cost. TeaCache first modulates the noisy inputs using the timestep embeddings to ensure their differences better approximating those of model outputs. TeaCache then introduces a rescaling strategy to refine the estimated differences and utilizes them to indicate output caching. Experiments show that TeaCache achieves up to 4.41x acceleration over Open-Sora-Plan with negligible (-0.07% Vbench score) degradation of visual quality.
title Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.19108