Adaptive Hybrid Caching for Efficient Text-to-Video Diffusion Model Acceleration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yuanxin, Diao, Lansong, Chen, Bujiao, Cheng, Shenggan, Qian, Zhengping, Yu, Wenyuan, Xiao, Nong, Lin, Wei, Du, Jiangsu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917294608744448
author Wei, Yuanxin
Diao, Lansong
Chen, Bujiao
Cheng, Shenggan
Qian, Zhengping
Yu, Wenyuan
Xiao, Nong
Lin, Wei
Du, Jiangsu
author_facet Wei, Yuanxin
Diao, Lansong
Chen, Bujiao
Cheng, Shenggan
Qian, Zhengping
Yu, Wenyuan
Xiao, Nong
Lin, Wei
Du, Jiangsu
contents Efficient video generation models are increasingly vital for multimedia synthetic content generation. Leveraging the Transformer architecture and the diffusion process, video DiT models have emerged as a dominant approach for high-quality video generation. However, their multi-step iterative denoising process incurs high computational cost and inference latency. Caching, a widely adopted optimization method in DiT models, leverages the redundancy in the diffusion process to skip computations in different granularities (e.g., step, cfg, block). Nevertheless, existing caching methods are limited to single-granularity strategies, struggling to balance generation quality and inference speed in a flexible manner. In this work, we propose MixCache, a training-free caching-based framework for efficient video DiT inference. It first distinguishes the interference and boundary between different caching strategies, and then introduces a context-aware cache triggering strategy to determine when caching should be enabled, along with an adaptive hybrid cache decision strategy for dynamically selecting the optimal caching granularity. Extensive experiments on diverse models demonstrate that, MixCache can significantly accelerate video generation (e.g., 1.94$\times$ speedup on Wan 14B, 1.97$\times$ speedup on HunyuanVideo) while delivering both superior generation quality and inference efficiency compared to baseline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12691
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Hybrid Caching for Efficient Text-to-Video Diffusion Model Acceleration
Wei, Yuanxin
Diao, Lansong
Chen, Bujiao
Cheng, Shenggan
Qian, Zhengping
Yu, Wenyuan
Xiao, Nong
Lin, Wei
Du, Jiangsu
Graphics
Computer Vision and Pattern Recognition
Machine Learning
Efficient video generation models are increasingly vital for multimedia synthetic content generation. Leveraging the Transformer architecture and the diffusion process, video DiT models have emerged as a dominant approach for high-quality video generation. However, their multi-step iterative denoising process incurs high computational cost and inference latency. Caching, a widely adopted optimization method in DiT models, leverages the redundancy in the diffusion process to skip computations in different granularities (e.g., step, cfg, block). Nevertheless, existing caching methods are limited to single-granularity strategies, struggling to balance generation quality and inference speed in a flexible manner. In this work, we propose MixCache, a training-free caching-based framework for efficient video DiT inference. It first distinguishes the interference and boundary between different caching strategies, and then introduces a context-aware cache triggering strategy to determine when caching should be enabled, along with an adaptive hybrid cache decision strategy for dynamically selecting the optimal caching granularity. Extensive experiments on diverse models demonstrate that, MixCache can significantly accelerate video generation (e.g., 1.94$\times$ speedup on Wan 14B, 1.97$\times$ speedup on HunyuanVideo) while delivering both superior generation quality and inference efficiency compared to baseline methods.
title Adaptive Hybrid Caching for Efficient Text-to-Video Diffusion Model Acceleration
topic Graphics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2508.12691