Goku: Flow Based Video Generative Foundation Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866916605624057856 |
|---|---|
| author | Chen, Shoufa Ge, Chongjian Zhang, Yuqi Zhang, Yida Zhu, Fengda Yang, Hao Hao, Hongxiang Wu, Hui Lai, Zhichao Hu, Yifei Lin, Ting-Che Zhang, Shilong Li, Fu Li, Chuan Wang, Xing Peng, Yanghua Sun, Peize Luo, Ping Jiang, Yi Yuan, Zehuan Peng, Bingyue Liu, Xiaobing |
| author_facet | Chen, Shoufa Ge, Chongjian Zhang, Yuqi Zhang, Yida Zhu, Fengda Yang, Hao Hao, Hongxiang Wu, Hui Lai, Zhichao Hu, Yifei Lin, Ting-Che Zhang, Shilong Li, Fu Li, Chuan Wang, Xing Peng, Yanghua Sun, Peize Luo, Ping Jiang, Yi Yuan, Zehuan Peng, Bingyue Liu, Xiaobing |
| contents | This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model architecture design, flow formulation, and advanced infrastructure for efficient and robust large-scale training. The Goku models demonstrate superior performance in both qualitative and quantitative evaluations, setting new benchmarks across major tasks. Specifically, Goku achieves 0.76 on GenEval and 83.65 on DPG-Bench for text-to-image generation, and 84.85 on VBench for text-to-video tasks. We believe that this work provides valuable insights and practical advancements for the research community in developing joint image-and-video generation models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_04896 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Goku: Flow Based Video Generative Foundation Models Chen, Shoufa Ge, Chongjian Zhang, Yuqi Zhang, Yida Zhu, Fengda Yang, Hao Hao, Hongxiang Wu, Hui Lai, Zhichao Hu, Yifei Lin, Ting-Che Zhang, Shilong Li, Fu Li, Chuan Wang, Xing Peng, Yanghua Sun, Peize Luo, Ping Jiang, Yi Yuan, Zehuan Peng, Bingyue Liu, Xiaobing Computer Vision and Pattern Recognition This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model architecture design, flow formulation, and advanced infrastructure for efficient and robust large-scale training. The Goku models demonstrate superior performance in both qualitative and quantitative evaluations, setting new benchmarks across major tasks. Specifically, Goku achieves 0.76 on GenEval and 83.65 on DPG-Bench for text-to-image generation, and 84.85 on VBench for text-to-video tasks. We believe that this work provides valuable insights and practical advancements for the research community in developing joint image-and-video generation models. |
| title | Goku: Flow Based Video Generative Foundation Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2502.04896 |