Goku: Flow Based Video Generative Foundation Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Shoufa, Ge, Chongjian, Zhang, Yuqi, Zhang, Yida, Zhu, Fengda, Yang, Hao, Hao, Hongxiang, Wu, Hui, Lai, Zhichao, Hu, Yifei, Lin, Ting-Che, Zhang, Shilong, Li, Fu, Li, Chuan, Wang, Xing, Peng, Yanghua, Sun, Peize, Luo, Ping, Jiang, Yi, Yuan, Zehuan, Peng, Bingyue, Liu, Xiaobing
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916605624057856
author Chen, Shoufa
Ge, Chongjian
Zhang, Yuqi
Zhang, Yida
Zhu, Fengda
Yang, Hao
Hao, Hongxiang
Wu, Hui
Lai, Zhichao
Hu, Yifei
Lin, Ting-Che
Zhang, Shilong
Li, Fu
Li, Chuan
Wang, Xing
Peng, Yanghua
Sun, Peize
Luo, Ping
Jiang, Yi
Yuan, Zehuan
Peng, Bingyue
Liu, Xiaobing
author_facet Chen, Shoufa
Ge, Chongjian
Zhang, Yuqi
Zhang, Yida
Zhu, Fengda
Yang, Hao
Hao, Hongxiang
Wu, Hui
Lai, Zhichao
Hu, Yifei
Lin, Ting-Che
Zhang, Shilong
Li, Fu
Li, Chuan
Wang, Xing
Peng, Yanghua
Sun, Peize
Luo, Ping
Jiang, Yi
Yuan, Zehuan
Peng, Bingyue
Liu, Xiaobing
contents This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model architecture design, flow formulation, and advanced infrastructure for efficient and robust large-scale training. The Goku models demonstrate superior performance in both qualitative and quantitative evaluations, setting new benchmarks across major tasks. Specifically, Goku achieves 0.76 on GenEval and 83.65 on DPG-Bench for text-to-image generation, and 84.85 on VBench for text-to-video tasks. We believe that this work provides valuable insights and practical advancements for the research community in developing joint image-and-video generation models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04896
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Goku: Flow Based Video Generative Foundation Models
Chen, Shoufa
Ge, Chongjian
Zhang, Yuqi
Zhang, Yida
Zhu, Fengda
Yang, Hao
Hao, Hongxiang
Wu, Hui
Lai, Zhichao
Hu, Yifei
Lin, Ting-Che
Zhang, Shilong
Li, Fu
Li, Chuan
Wang, Xing
Peng, Yanghua
Sun, Peize
Luo, Ping
Jiang, Yi
Yuan, Zehuan
Peng, Bingyue
Liu, Xiaobing
Computer Vision and Pattern Recognition
This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model architecture design, flow formulation, and advanced infrastructure for efficient and robust large-scale training. The Goku models demonstrate superior performance in both qualitative and quantitative evaluations, setting new benchmarks across major tasks. Specifically, Goku achieves 0.76 on GenEval and 83.65 on DPG-Bench for text-to-image generation, and 84.85 on VBench for text-to-video tasks. We believe that this work provides valuable insights and practical advancements for the research community in developing joint image-and-video generation models.
title Goku: Flow Based Video Generative Foundation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.04896