VideoPoet: A Large Language Model for Zero-Shot Video Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kondratyuk, Dan, Yu, Lijun, Gu, Xiuye, Lezama, José, Huang, Jonathan, Schindler, Grant, Hornung, Rachel, Birodkar, Vighnesh, Yan, Jimmy, Chiu, Ming-Chang, Somandepalli, Krishna, Akbari, Hassan, Alon, Yair, Cheng, Yong, Dillon, Josh, Gupta, Agrim, Hahn, Meera, Hauth, Anja, Hendon, David, Martinez, Alonso, Minnen, David, Sirotenko, Mikhail, Sohn, Kihyuk, Yang, Xuan, Adam, Hartwig, Yang, Ming-Hsuan, Essa, Irfan, Wang, Huisheng, Ross, David A., Seybold, Bryan, Jiang, Lu
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929372304244736
author Kondratyuk, Dan
Yu, Lijun
Gu, Xiuye
Lezama, José
Huang, Jonathan
Schindler, Grant
Hornung, Rachel
Birodkar, Vighnesh
Yan, Jimmy
Chiu, Ming-Chang
Somandepalli, Krishna
Akbari, Hassan
Alon, Yair
Cheng, Yong
Dillon, Josh
Gupta, Agrim
Hahn, Meera
Hauth, Anja
Hendon, David
Martinez, Alonso
Minnen, David
Sirotenko, Mikhail
Sohn, Kihyuk
Yang, Xuan
Adam, Hartwig
Yang, Ming-Hsuan
Essa, Irfan
Wang, Huisheng
Ross, David A.
Seybold, Bryan
Jiang, Lu
author_facet Kondratyuk, Dan
Yu, Lijun
Gu, Xiuye
Lezama, José
Huang, Jonathan
Schindler, Grant
Hornung, Rachel
Birodkar, Vighnesh
Yan, Jimmy
Chiu, Ming-Chang
Somandepalli, Krishna
Akbari, Hassan
Alon, Yair
Cheng, Yong
Dillon, Josh
Gupta, Agrim
Hahn, Meera
Hauth, Anja
Hendon, David
Martinez, Alonso
Minnen, David
Sirotenko, Mikhail
Sohn, Kihyuk
Yang, Xuan
Adam, Hartwig
Yang, Ming-Hsuan
Essa, Irfan
Wang, Huisheng
Ross, David A.
Seybold, Bryan
Jiang, Lu
contents We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs -- including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model's state-of-the-art capabilities in zero-shot video generation, specifically highlighting VideoPoet's ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/
format Preprint
id arxiv_https___arxiv_org_abs_2312_14125
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VideoPoet: A Large Language Model for Zero-Shot Video Generation
Kondratyuk, Dan
Yu, Lijun
Gu, Xiuye
Lezama, José
Huang, Jonathan
Schindler, Grant
Hornung, Rachel
Birodkar, Vighnesh
Yan, Jimmy
Chiu, Ming-Chang
Somandepalli, Krishna
Akbari, Hassan
Alon, Yair
Cheng, Yong
Dillon, Josh
Gupta, Agrim
Hahn, Meera
Hauth, Anja
Hendon, David
Martinez, Alonso
Minnen, David
Sirotenko, Mikhail
Sohn, Kihyuk
Yang, Xuan
Adam, Hartwig
Yang, Ming-Hsuan
Essa, Irfan
Wang, Huisheng
Ross, David A.
Seybold, Bryan
Jiang, Lu
Computer Vision and Pattern Recognition
Artificial Intelligence
We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs -- including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model's state-of-the-art capabilities in zero-shot video generation, specifically highlighting VideoPoet's ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/
title VideoPoet: A Large Language Model for Zero-Shot Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2312.14125