CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Qi, Xingqun, Zhang, Hengyuan, Wang, Yatian, Pan, Jiahao, Liu, Chen, Li, Peng, Chi, Xiaowei, Li, Mengfei, Xue, Wei, Zhang, Shanghang, Luo, Wenhan, Liu, Qifeng, Guo, Yike
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917998022885376
author Qi, Xingqun
Zhang, Hengyuan
Wang, Yatian
Pan, Jiahao
Liu, Chen
Li, Peng
Chi, Xiaowei
Li, Mengfei
Xue, Wei
Zhang, Shanghang
Luo, Wenhan
Liu, Qifeng
Guo, Yike
author_facet Qi, Xingqun
Zhang, Hengyuan
Wang, Yatian
Pan, Jiahao
Liu, Chen
Li, Peng
Chi, Xiaowei
Li, Mengfei
Xue, Wei
Zhang, Shanghang
Luo, Wenhan
Liu, Qifeng
Guo, Yike
contents Deriving co-speech 3D gestures has seen tremendous progress in virtual avatar animation. Yet, the existing methods often produce stiff and unreasonable gestures with unseen human speech inputs due to the limited 3D speech-gesture data. In this paper, we propose CoCoGesture, a novel framework enabling vivid and diverse gesture synthesis from unseen human speech prompts. Our key insight is built upon the custom-designed pretrain-fintune training paradigm. At the pretraining stage, we aim to formulate a large generalizable gesture diffusion model by learning the abundant postures manifold. Therefore, to alleviate the scarcity of 3D data, we first construct a large-scale co-speech 3D gesture dataset containing more than 40M meshed posture instances across 4.3K speakers, dubbed GES-X. Then, we scale up the large unconditional diffusion model to 1B parameters and pre-train it to be our gesture experts. At the finetune stage, we present the audio ControlNet that incorporates the human voice as condition prompts to guide the gesture generation. Here, we construct the audio ControlNet through a trainable copy of our pre-trained diffusion model. Moreover, we design a novel Mixture-of-Gesture-Experts (MoGE) block to adaptively fuse the audio embedding from the human speech and the gesture features from the pre-trained gesture experts with a routing mechanism. Such an effective manner ensures audio embedding is temporal coordinated with motion features while preserving the vivid and diverse gesture generation. Extensive experiments demonstrate that our proposed CoCoGesture outperforms the state-of-the-art methods on the zero-shot speech-to-gesture generation. The dataset will be publicly available at: https://mattie-e.github.io/GES-X/
format Preprint
id arxiv_https___arxiv_org_abs_2405_16874
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild
Qi, Xingqun
Zhang, Hengyuan
Wang, Yatian
Pan, Jiahao
Liu, Chen
Li, Peng
Chi, Xiaowei
Li, Mengfei
Xue, Wei
Zhang, Shanghang
Luo, Wenhan
Liu, Qifeng
Guo, Yike
Computer Vision and Pattern Recognition
Deriving co-speech 3D gestures has seen tremendous progress in virtual avatar animation. Yet, the existing methods often produce stiff and unreasonable gestures with unseen human speech inputs due to the limited 3D speech-gesture data. In this paper, we propose CoCoGesture, a novel framework enabling vivid and diverse gesture synthesis from unseen human speech prompts. Our key insight is built upon the custom-designed pretrain-fintune training paradigm. At the pretraining stage, we aim to formulate a large generalizable gesture diffusion model by learning the abundant postures manifold. Therefore, to alleviate the scarcity of 3D data, we first construct a large-scale co-speech 3D gesture dataset containing more than 40M meshed posture instances across 4.3K speakers, dubbed GES-X. Then, we scale up the large unconditional diffusion model to 1B parameters and pre-train it to be our gesture experts. At the finetune stage, we present the audio ControlNet that incorporates the human voice as condition prompts to guide the gesture generation. Here, we construct the audio ControlNet through a trainable copy of our pre-trained diffusion model. Moreover, we design a novel Mixture-of-Gesture-Experts (MoGE) block to adaptively fuse the audio embedding from the human speech and the gesture features from the pre-trained gesture experts with a routing mechanism. Such an effective manner ensures audio embedding is temporal coordinated with motion features while preserving the vivid and diverse gesture generation. Extensive experiments demonstrate that our proposed CoCoGesture outperforms the state-of-the-art methods on the zero-shot speech-to-gesture generation. The dataset will be publicly available at: https://mattie-e.github.io/GES-X/
title CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.16874