DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Junming, Liu, Yunfei, Wang, Jianan, Zeng, Ailing, Li, Yu, Chen, Qifeng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909162845241344
author Chen, Junming
Liu, Yunfei
Wang, Jianan
Zeng, Ailing
Li, Yu
Chen, Qifeng
author_facet Chen, Junming
Liu, Yunfei
Wang, Jianan
Zeng, Ailing
Li, Yu
Chen, Qifeng
contents We propose DiffSHEG, a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation with arbitrary length. While previous works focused on co-speech gesture or expression generation individually, the joint generation of synchronized expressions and gestures remains barely explored. To address this, our diffusion-based co-speech motion generation transformer enables uni-directional information flow from expression to gesture, facilitating improved matching of joint expression-gesture distributions. Furthermore, we introduce an outpainting-based sampling strategy for arbitrary long sequence generation in diffusion models, offering flexibility and computational efficiency. Our method provides a practical solution that produces high-quality synchronized expression and gesture generation driven by speech. Evaluated on two public datasets, our approach achieves state-of-the-art performance both quantitatively and qualitatively. Additionally, a user study confirms the superiority of DiffSHEG over prior approaches. By enabling the real-time generation of expressive and synchronized motions, DiffSHEG showcases its potential for various applications in the development of digital humans and embodied agents.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04747
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation
Chen, Junming
Liu, Yunfei
Wang, Jianan
Zeng, Ailing
Li, Yu
Chen, Qifeng
Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Graphics
Audio and Speech Processing
We propose DiffSHEG, a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation with arbitrary length. While previous works focused on co-speech gesture or expression generation individually, the joint generation of synchronized expressions and gestures remains barely explored. To address this, our diffusion-based co-speech motion generation transformer enables uni-directional information flow from expression to gesture, facilitating improved matching of joint expression-gesture distributions. Furthermore, we introduce an outpainting-based sampling strategy for arbitrary long sequence generation in diffusion models, offering flexibility and computational efficiency. Our method provides a practical solution that produces high-quality synchronized expression and gesture generation driven by speech. Evaluated on two public datasets, our approach achieves state-of-the-art performance both quantitatively and qualitatively. Additionally, a user study confirms the superiority of DiffSHEG over prior approaches. By enabling the real-time generation of expressive and synchronized motions, DiffSHEG showcases its potential for various applications in the development of digital humans and embodied agents.
title DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation
topic Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Graphics
Audio and Speech Processing
url https://arxiv.org/abs/2401.04747