Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909825587216384 |
|---|---|
| author | Zheng, Shikang Chen, Guantao Zhou, Qinming Lin, Yuqi He, Lixuan Zou, Chang Cai, Peiliang Liu, Jiacheng Zhang, Linfeng |
| author_facet | Zheng, Shikang Chen, Guantao Zhou, Qinming Lin, Yuqi He, Lixuan Zou, Chang Cai, Peiliang Liu, Jiacheng Zhang, Linfeng |
| contents | Diffusion Transformers offer state-of-the-art fidelity in image and video synthesis, but their iterative sampling process remains a major bottleneck due to the high cost of transformer forward passes at each timestep. To mitigate this, feature caching has emerged as a training-free acceleration technique that reuses or forecasts hidden representations. However, existing methods often apply a uniform caching strategy across all feature dimensions, ignoring their heterogeneous dynamic behaviors. Therefore, we adopt a new perspective by modeling hidden feature evolution as a mixture of ODEs across dimensions, and introduce HyCa, a Hybrid ODE solver inspired caching framework that applies dimension-wise caching strategies. HyCa achieves near-lossless acceleration across diverse domains and models, including 5.55 times speedup on FLUX, 5.56 times speedup on HunyuanVideo, 6.24 times speedup on Qwen-Image and Qwen-Image-Edit without retraining. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_04188 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers Zheng, Shikang Chen, Guantao Zhou, Qinming Lin, Yuqi He, Lixuan Zou, Chang Cai, Peiliang Liu, Jiacheng Zhang, Linfeng Computer Vision and Pattern Recognition Diffusion Transformers offer state-of-the-art fidelity in image and video synthesis, but their iterative sampling process remains a major bottleneck due to the high cost of transformer forward passes at each timestep. To mitigate this, feature caching has emerged as a training-free acceleration technique that reuses or forecasts hidden representations. However, existing methods often apply a uniform caching strategy across all feature dimensions, ignoring their heterogeneous dynamic behaviors. Therefore, we adopt a new perspective by modeling hidden feature evolution as a mixture of ODEs across dimensions, and introduce HyCa, a Hybrid ODE solver inspired caching framework that applies dimension-wise caching strategies. HyCa achieves near-lossless acceleration across diverse domains and models, including 5.55 times speedup on FLUX, 5.56 times speedup on HunyuanVideo, 6.24 times speedup on Qwen-Image and Qwen-Image-Edit without retraining. |
| title | Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.04188 |