Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion Generations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Ruoxi, Pi, Huaijin, Shen, Zehong, Shuai, Qing, Hu, Zechen, Wang, Zhumei, Dong, Yajiao, Hu, Ruizhen, Komura, Taku, Peng, Sida, Zhou, Xiaowei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917510384713728
author Guo, Ruoxi
Pi, Huaijin
Shen, Zehong
Shuai, Qing
Hu, Zechen
Wang, Zhumei
Dong, Yajiao
Hu, Ruizhen
Komura, Taku
Peng, Sida
Zhou, Xiaowei
author_facet Guo, Ruoxi
Pi, Huaijin
Shen, Zehong
Shuai, Qing
Hu, Zechen
Wang, Zhumei
Dong, Yajiao
Hu, Ruizhen
Komura, Taku
Peng, Sida
Zhou, Xiaowei
contents Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods often rely on 3D motion capture data, which requires special setups, resulting in high costs for data acquisition, ultimately limiting the diversity and scope of human motion. In contrast, 2D human videos offer a vast and accessible source of motion data, covering a wider range of styles and activities. In this paper, we explore the use of 2D human motion extracted from videos as an alternative data source to improve text-driven 3D motion generation. Our approach introduces a novel framework that disentangles local joint motion from global movements, enabling efficient learning of local motion priors from 2D data. We first train a single-view 2D local motion generator on a large dataset of text-2D motion pairs. Then we fine-tune the generator with 3D data, transforming it into a multi-view generator that predicts view-consistent local joint motion and root dynamics. Evaluations on the well-acknowledged dataset and novel text prompts demonstrate that our method can efficiently utilize 2D data, supporting a wider range of realistic 3D human motion generation. Our code is publicly available at https://zju3dv.github.io/Motion-2-to-3/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13111
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion Generations
Guo, Ruoxi
Pi, Huaijin
Shen, Zehong
Shuai, Qing
Hu, Zechen
Wang, Zhumei
Dong, Yajiao
Hu, Ruizhen
Komura, Taku
Peng, Sida
Zhou, Xiaowei
Computer Vision and Pattern Recognition
Graphics
Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods often rely on 3D motion capture data, which requires special setups, resulting in high costs for data acquisition, ultimately limiting the diversity and scope of human motion. In contrast, 2D human videos offer a vast and accessible source of motion data, covering a wider range of styles and activities. In this paper, we explore the use of 2D human motion extracted from videos as an alternative data source to improve text-driven 3D motion generation. Our approach introduces a novel framework that disentangles local joint motion from global movements, enabling efficient learning of local motion priors from 2D data. We first train a single-view 2D local motion generator on a large dataset of text-2D motion pairs. Then we fine-tune the generator with 3D data, transforming it into a multi-view generator that predicts view-consistent local joint motion and root dynamics. Evaluations on the well-acknowledged dataset and novel text prompts demonstrate that our method can efficiently utilize 2D data, supporting a wider range of realistic 3D human motion generation. Our code is publicly available at https://zju3dv.github.io/Motion-2-to-3/.
title Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion Generations
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2412.13111