Fast and Memory-Efficient Video Diffusion Using Streamlined Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhan, Zheng, Wu, Yushu, Gong, Yifan, Meng, Zichong, Kong, Zhenglun, Yang, Changdi, Yuan, Geng, Zhao, Pu, Niu, Wei, Wang, Yanzhi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917827357704192
author Zhan, Zheng
Wu, Yushu
Gong, Yifan
Meng, Zichong
Kong, Zhenglun
Yang, Changdi
Yuan, Geng
Zhao, Pu
Niu, Wei
Wang, Yanzhi
author_facet Zhan, Zheng
Wu, Yushu
Gong, Yifan
Meng, Zichong
Kong, Zhenglun
Yang, Changdi
Yuan, Geng
Zhao, Pu
Niu, Wei
Wang, Yanzhi
contents The rapid progress in artificial intelligence-generated content (AIGC), especially with diffusion models, has significantly advanced development of high-quality video generation. However, current video diffusion models exhibit demanding computational requirements and high peak memory usage, especially for generating longer and higher-resolution videos. These limitations greatly hinder the practical application of video diffusion models on standard hardware platforms. To tackle this issue, we present a novel, training-free framework named Streamlined Inference, which leverages the temporal and spatial properties of video diffusion models. Our approach integrates three core components: Feature Slicer, Operator Grouping, and Step Rehash. Specifically, Feature Slicer effectively partitions input features into sub-features and Operator Grouping processes each sub-feature with a group of consecutive operators, resulting in significant memory reduction without sacrificing the quality or speed. Step Rehash further exploits the similarity between adjacent steps in diffusion, and accelerates inference through skipping unnecessary steps. Extensive experiments demonstrate that our approach significantly reduces peak memory and computational overhead, making it feasible to generate high-quality videos on a single consumer GPU (e.g., reducing peak memory of AnimateDiff from 42GB to 11GB, featuring faster inference on 2080Ti).
format Preprint
id arxiv_https___arxiv_org_abs_2411_01171
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fast and Memory-Efficient Video Diffusion Using Streamlined Inference
Zhan, Zheng
Wu, Yushu
Gong, Yifan
Meng, Zichong
Kong, Zhenglun
Yang, Changdi
Yuan, Geng
Zhao, Pu
Niu, Wei
Wang, Yanzhi
Computer Vision and Pattern Recognition
Artificial Intelligence
The rapid progress in artificial intelligence-generated content (AIGC), especially with diffusion models, has significantly advanced development of high-quality video generation. However, current video diffusion models exhibit demanding computational requirements and high peak memory usage, especially for generating longer and higher-resolution videos. These limitations greatly hinder the practical application of video diffusion models on standard hardware platforms. To tackle this issue, we present a novel, training-free framework named Streamlined Inference, which leverages the temporal and spatial properties of video diffusion models. Our approach integrates three core components: Feature Slicer, Operator Grouping, and Step Rehash. Specifically, Feature Slicer effectively partitions input features into sub-features and Operator Grouping processes each sub-feature with a group of consecutive operators, resulting in significant memory reduction without sacrificing the quality or speed. Step Rehash further exploits the similarity between adjacent steps in diffusion, and accelerates inference through skipping unnecessary steps. Extensive experiments demonstrate that our approach significantly reduces peak memory and computational overhead, making it feasible to generate high-quality videos on a single consumer GPU (e.g., reducing peak memory of AnimateDiff from 42GB to 11GB, featuring faster inference on 2080Ti).
title Fast and Memory-Efficient Video Diffusion Using Streamlined Inference
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.01171