SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhengang, Kang, Yan, Liu, Yuchen, Liu, Difan, Hinz, Tobias, Liu, Feng, Wang, Yanzhi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917681731469312
author Li, Zhengang
Kang, Yan
Liu, Yuchen
Liu, Difan
Hinz, Tobias
Liu, Feng
Wang, Yanzhi
author_facet Li, Zhengang
Kang, Yan
Liu, Yuchen
Liu, Difan
Hinz, Tobias
Liu, Feng
Wang, Yanzhi
contents While AI-generated content has garnered significant attention, achieving photo-realistic video synthesis remains a formidable challenge. Despite the promising advances in diffusion models for video generation quality, the complex model architecture and substantial computational demands for both training and inference create a significant gap between these models and real-world applications. This paper presents SNED, a superposition network architecture search method for efficient video diffusion model. Our method employs a supernet training paradigm that targets various model cost and resolution options using a weight-sharing method. Moreover, we propose the supernet training sampling warm-up for fast training optimization. To showcase the flexibility of our method, we conduct experiments involving both pixel-space and latent-space video diffusion models. The results demonstrate that our framework consistently produces comparable results across different model options with high efficiency. According to the experiment for the pixel-space video diffusion model, we can achieve consistent video generation results simultaneously across 64 x 64 to 256 x 256 resolutions with a large range of model sizes from 640M to 1.6B number of parameters for pixel-space video diffusion models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00195
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model
Li, Zhengang
Kang, Yan
Liu, Yuchen
Liu, Difan
Hinz, Tobias
Liu, Feng
Wang, Yanzhi
Computer Vision and Pattern Recognition
Artificial Intelligence
While AI-generated content has garnered significant attention, achieving photo-realistic video synthesis remains a formidable challenge. Despite the promising advances in diffusion models for video generation quality, the complex model architecture and substantial computational demands for both training and inference create a significant gap between these models and real-world applications. This paper presents SNED, a superposition network architecture search method for efficient video diffusion model. Our method employs a supernet training paradigm that targets various model cost and resolution options using a weight-sharing method. Moreover, we propose the supernet training sampling warm-up for fast training optimization. To showcase the flexibility of our method, we conduct experiments involving both pixel-space and latent-space video diffusion models. The results demonstrate that our framework consistently produces comparable results across different model options with high efficiency. According to the experiment for the pixel-space video diffusion model, we can achieve consistent video generation results simultaneously across 64 x 64 to 256 x 256 resolutions with a large range of model sizes from 640M to 1.6B number of parameters for pixel-space video diffusion models.
title SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2406.00195