Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Shikang, Chen, Guantao, Zhou, Qinming, Lin, Yuqi, He, Lixuan, Zou, Chang, Cai, Peiliang, Liu, Jiacheng, Zhang, Linfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909825587216384
author Zheng, Shikang
Chen, Guantao
Zhou, Qinming
Lin, Yuqi
He, Lixuan
Zou, Chang
Cai, Peiliang
Liu, Jiacheng
Zhang, Linfeng
author_facet Zheng, Shikang
Chen, Guantao
Zhou, Qinming
Lin, Yuqi
He, Lixuan
Zou, Chang
Cai, Peiliang
Liu, Jiacheng
Zhang, Linfeng
contents Diffusion Transformers offer state-of-the-art fidelity in image and video synthesis, but their iterative sampling process remains a major bottleneck due to the high cost of transformer forward passes at each timestep. To mitigate this, feature caching has emerged as a training-free acceleration technique that reuses or forecasts hidden representations. However, existing methods often apply a uniform caching strategy across all feature dimensions, ignoring their heterogeneous dynamic behaviors. Therefore, we adopt a new perspective by modeling hidden feature evolution as a mixture of ODEs across dimensions, and introduce HyCa, a Hybrid ODE solver inspired caching framework that applies dimension-wise caching strategies. HyCa achieves near-lossless acceleration across diverse domains and models, including 5.55 times speedup on FLUX, 5.56 times speedup on HunyuanVideo, 6.24 times speedup on Qwen-Image and Qwen-Image-Edit without retraining.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
Zheng, Shikang
Chen, Guantao
Zhou, Qinming
Lin, Yuqi
He, Lixuan
Zou, Chang
Cai, Peiliang
Liu, Jiacheng
Zhang, Linfeng
Computer Vision and Pattern Recognition
Diffusion Transformers offer state-of-the-art fidelity in image and video synthesis, but their iterative sampling process remains a major bottleneck due to the high cost of transformer forward passes at each timestep. To mitigate this, feature caching has emerged as a training-free acceleration technique that reuses or forecasts hidden representations. However, existing methods often apply a uniform caching strategy across all feature dimensions, ignoring their heterogeneous dynamic behaviors. Therefore, we adopt a new perspective by modeling hidden feature evolution as a mixture of ODEs across dimensions, and introduce HyCa, a Hybrid ODE solver inspired caching framework that applies dimension-wise caching strategies. HyCa achieves near-lossless acceleration across diverse domains and models, including 5.55 times speedup on FLUX, 5.56 times speedup on HunyuanVideo, 6.24 times speedup on Qwen-Image and Qwen-Image-Edit without retraining.
title Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.04188