Pyramidal Flow Matching for Efficient Video Generative Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jin, Yang, Sun, Zhicheng, Li, Ningyuan, Xu, Kun, Jiang, Hao, Zhuang, Nan, Huang, Quzhe, Song, Yang, Mu, Yadong, Lin, Zhouchen
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909538109620224
author Jin, Yang
Sun, Zhicheng
Li, Ningyuan
Xu, Kun
Xu, Kun
Jiang, Hao
Zhuang, Nan
Huang, Quzhe
Song, Yang
Mu, Yadong
Lin, Zhouchen
author_facet Jin, Yang
Sun, Zhicheng
Li, Ningyuan
Xu, Kun
Xu, Kun
Jiang, Hao
Zhuang, Nan
Huang, Quzhe
Song, Yang
Mu, Yadong
Lin, Zhouchen
contents Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational demands, the separate optimization of each sub-stage hinders knowledge sharing and sacrifices flexibility. This work introduces a unified pyramidal flow matching algorithm. It reinterprets the original denoising trajectory as a series of pyramid stages, where only the final stage operates at the full resolution, thereby enabling more efficient video generative modeling. Through our sophisticated design, the flows of different pyramid stages can be interlinked to maintain continuity. Moreover, we craft autoregressive video generation with a temporal pyramid to compress the full-resolution history. The entire framework can be optimized in an end-to-end manner and with a single unified Diffusion Transformer (DiT). Extensive experiments demonstrate that our method supports generating high-quality 5-second (up to 10-second) videos at 768p resolution and 24 FPS within 20.7k A100 GPU training hours. All code and models are open-sourced at https://pyramid-flow.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05954
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pyramidal Flow Matching for Efficient Video Generative Modeling
Jin, Yang
Sun, Zhicheng
Li, Ningyuan
Xu, Kun
Xu, Kun
Jiang, Hao
Zhuang, Nan
Huang, Quzhe
Song, Yang
Mu, Yadong
Lin, Zhouchen
Computer Vision and Pattern Recognition
Machine Learning
Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational demands, the separate optimization of each sub-stage hinders knowledge sharing and sacrifices flexibility. This work introduces a unified pyramidal flow matching algorithm. It reinterprets the original denoising trajectory as a series of pyramid stages, where only the final stage operates at the full resolution, thereby enabling more efficient video generative modeling. Through our sophisticated design, the flows of different pyramid stages can be interlinked to maintain continuity. Moreover, we craft autoregressive video generation with a temporal pyramid to compress the full-resolution history. The entire framework can be optimized in an end-to-end manner and with a single unified Diffusion Transformer (DiT). Extensive experiments demonstrate that our method supports generating high-quality 5-second (up to 10-second) videos at 768p resolution and 24 FPS within 20.7k A100 GPU training hours. All code and models are open-sourced at https://pyramid-flow.github.io.
title Pyramidal Flow Matching for Efficient Video Generative Modeling
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.05954