MagCache: Fast Video Generation with Magnitude-Aware Cache

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ma, Zehong, Wei, Longhui, Wang, Feng, Zhang, Shiliang, Tian, Qi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908628928167936
author Ma, Zehong
Wei, Longhui
Wang, Feng
Zhang, Shiliang
Tian, Qi
author_facet Ma, Zehong
Wei, Longhui
Wang, Feng
Zhang, Shiliang
Tian, Qi
contents Existing acceleration techniques for video diffusion models often rely on uniform heuristics or time-embedding variants to skip timesteps and reuse cached features. These approaches typically require extensive calibration with curated prompts and risk inconsistent outputs due to prompt-specific overfitting. In this paper, we introduce a novel and robust discovery: a unified magnitude law observed across different models and prompts. Specifically, the magnitude ratio of successive residual outputs decreases monotonically, steadily in most timesteps while rapidly in the last several steps. Leveraging this insight, we introduce a Magnitude-aware Cache (MagCache) that adaptively skips unimportant timesteps using an error modeling mechanism and adaptive caching strategy. Unlike existing methods requiring dozens of curated samples for calibration, MagCache only requires a single sample for calibration. Experimental results show that MagCache achieves 2.10x-2.68x speedups on Open-Sora, CogVideoX, Wan 2.1, and HunyuanVideo, while preserving superior visual fidelity. It significantly outperforms existing methods in LPIPS, SSIM, and PSNR, under similar computational budgets.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09045
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MagCache: Fast Video Generation with Magnitude-Aware Cache
Ma, Zehong
Wei, Longhui
Wang, Feng
Zhang, Shiliang
Tian, Qi
Computer Vision and Pattern Recognition
Existing acceleration techniques for video diffusion models often rely on uniform heuristics or time-embedding variants to skip timesteps and reuse cached features. These approaches typically require extensive calibration with curated prompts and risk inconsistent outputs due to prompt-specific overfitting. In this paper, we introduce a novel and robust discovery: a unified magnitude law observed across different models and prompts. Specifically, the magnitude ratio of successive residual outputs decreases monotonically, steadily in most timesteps while rapidly in the last several steps. Leveraging this insight, we introduce a Magnitude-aware Cache (MagCache) that adaptively skips unimportant timesteps using an error modeling mechanism and adaptive caching strategy. Unlike existing methods requiring dozens of curated samples for calibration, MagCache only requires a single sample for calibration. Experimental results show that MagCache achieves 2.10x-2.68x speedups on Open-Sora, CogVideoX, Wan 2.1, and HunyuanVideo, while preserving superior visual fidelity. It significantly outperforms existing methods in LPIPS, SSIM, and PSNR, under similar computational budgets.
title MagCache: Fast Video Generation with Magnitude-Aware Cache
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.09045