Pyramidal Flow Matching for Efficient Video Generative Modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909538109620224 |
|---|---|
| author | Jin, Yang Sun, Zhicheng Li, Ningyuan Xu, Kun Xu, Kun Jiang, Hao Zhuang, Nan Huang, Quzhe Song, Yang Mu, Yadong Lin, Zhouchen |
| author_facet | Jin, Yang Sun, Zhicheng Li, Ningyuan Xu, Kun Xu, Kun Jiang, Hao Zhuang, Nan Huang, Quzhe Song, Yang Mu, Yadong Lin, Zhouchen |
| contents | Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational demands, the separate optimization of each sub-stage hinders knowledge sharing and sacrifices flexibility. This work introduces a unified pyramidal flow matching algorithm. It reinterprets the original denoising trajectory as a series of pyramid stages, where only the final stage operates at the full resolution, thereby enabling more efficient video generative modeling. Through our sophisticated design, the flows of different pyramid stages can be interlinked to maintain continuity. Moreover, we craft autoregressive video generation with a temporal pyramid to compress the full-resolution history. The entire framework can be optimized in an end-to-end manner and with a single unified Diffusion Transformer (DiT). Extensive experiments demonstrate that our method supports generating high-quality 5-second (up to 10-second) videos at 768p resolution and 24 FPS within 20.7k A100 GPU training hours. All code and models are open-sourced at https://pyramid-flow.github.io. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_05954 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Pyramidal Flow Matching for Efficient Video Generative Modeling Jin, Yang Sun, Zhicheng Li, Ningyuan Xu, Kun Xu, Kun Jiang, Hao Zhuang, Nan Huang, Quzhe Song, Yang Mu, Yadong Lin, Zhouchen Computer Vision and Pattern Recognition Machine Learning Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational demands, the separate optimization of each sub-stage hinders knowledge sharing and sacrifices flexibility. This work introduces a unified pyramidal flow matching algorithm. It reinterprets the original denoising trajectory as a series of pyramid stages, where only the final stage operates at the full resolution, thereby enabling more efficient video generative modeling. Through our sophisticated design, the flows of different pyramid stages can be interlinked to maintain continuity. Moreover, we craft autoregressive video generation with a temporal pyramid to compress the full-resolution history. The entire framework can be optimized in an end-to-end manner and with a single unified Diffusion Transformer (DiT). Extensive experiments demonstrate that our method supports generating high-quality 5-second (up to 10-second) videos at 768p resolution and 24 FPS within 20.7k A100 GPU training hours. All code and models are open-sourced at https://pyramid-flow.github.io. |
| title | Pyramidal Flow Matching for Efficient Video Generative Modeling |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2410.05954 |