OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Hui, Xu, Mingwang, Zhan, Yun, Mu, Shan, Li, Jiaye, Cheng, Kaihui, Chen, Yuxuan, Chen, Tan, Ye, Mao, Wang, Jingdong, Zhu, Siyu
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910771944882176
author Li, Hui
Xu, Mingwang
Zhan, Yun
Mu, Shan
Li, Jiaye
Cheng, Kaihui
Chen, Yuxuan
Chen, Tan
Ye, Mao
Wang, Jingdong
Zhu, Siyu
author_facet Li, Hui
Xu, Mingwang
Zhan, Yun
Mu, Shan
Li, Jiaye
Cheng, Kaihui
Chen, Yuxuan
Chen, Tan
Ye, Mao
Wang, Jingdong
Zhu, Siyu
contents Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality, human-centric video datasets presents a challenge to progress in this field. To bridge this gap, we introduce OpenHumanVid, a large-scale and high-quality human-centric video dataset characterized by precise and detailed captions that encompass both human appearance and motion states, along with supplementary human motion conditions, including skeleton sequences and speech audio. To validate the efficacy of this dataset and the associated training strategies, we propose an extension of existing classical diffusion transformer architectures and conduct further pretraining of our models on the proposed dataset. Our findings yield two critical insights: First, the incorporation of a large-scale, high-quality dataset substantially enhances evaluation metrics for generated human videos while preserving performance in general video generation tasks. Second, the effective alignment of text with human appearance, human motion, and facial motion is essential for producing high-quality video outputs. Based on these insights and corresponding methodologies, the straightforward extended network trained on the proposed dataset demonstrates an obvious improvement in the generation of human-centric videos. Project page https://fudan-generative-vision.github.io/OpenHumanVid
format Preprint
id arxiv_https___arxiv_org_abs_2412_00115
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
Li, Hui
Xu, Mingwang
Zhan, Yun
Mu, Shan
Li, Jiaye
Cheng, Kaihui
Chen, Yuxuan
Chen, Tan
Ye, Mao
Wang, Jingdong
Zhu, Siyu
Computer Vision and Pattern Recognition
Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality, human-centric video datasets presents a challenge to progress in this field. To bridge this gap, we introduce OpenHumanVid, a large-scale and high-quality human-centric video dataset characterized by precise and detailed captions that encompass both human appearance and motion states, along with supplementary human motion conditions, including skeleton sequences and speech audio. To validate the efficacy of this dataset and the associated training strategies, we propose an extension of existing classical diffusion transformer architectures and conduct further pretraining of our models on the proposed dataset. Our findings yield two critical insights: First, the incorporation of a large-scale, high-quality dataset substantially enhances evaluation metrics for generated human videos while preserving performance in general video generation tasks. Second, the effective alignment of text with human appearance, human motion, and facial motion is essential for producing high-quality video outputs. Based on these insights and corresponding methodologies, the straightforward extended network trained on the proposed dataset demonstrates an obvious improvement in the generation of human-centric videos. Project page https://fudan-generative-vision.github.io/OpenHumanVid
title OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00115