Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Xiang, Pang, Youxin, Zhao, Xiaochen, Xu, Chao, Wang, Lizhen, Xiao, Hongjiang, Yan, Shi, Zhang, Hongwen, Liu, Yebin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917297551048704
author Deng, Xiang
Pang, Youxin
Zhao, Xiaochen
Xu, Chao
Wang, Lizhen
Xiao, Hongjiang
Yan, Shi
Zhang, Hongwen
Liu, Yebin
author_facet Deng, Xiang
Pang, Youxin
Zhao, Xiaochen
Xu, Chao
Wang, Lizhen
Xiao, Hongjiang
Yan, Shi
Zhang, Hongwen
Liu, Yebin
contents This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-realistic quality, and continuous viewpoint control. The process follows a two-stage approach. In the first stage, the system maps audio input to high-fidelity motion sequences, encompassing upper-body gestures and facial expressions. To enrich motion diversity and authenticity, large language model (LLM) priors are integrated with text-aligned semantic audio features, leveraging LLMs' cross-modal generalization power to enhance motion quality. In the second stage, we improve diffusion-based video generation models by incorporating a prior-guided Mixture-of-Experts (MoE) mechanism: a view-guided MoE focuses on view-specific attributes, while a mask-guided MoE enhances region-based rendering stability. Additionally, a mask prediction module is devised to derive human masks from motion data, enhancing the stability and accuracy of masks and enabling mask guiding during inference. We also introduce a comprehensive human video dataset with 2,203 identities, covering diverse body gestures and detailed annotations, facilitating broad generalization. The code, data, and pre-trained models will be released for research purposes.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23836
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts
Deng, Xiang
Pang, Youxin
Zhao, Xiaochen
Xu, Chao
Wang, Lizhen
Xiao, Hongjiang
Yan, Shi
Zhang, Hongwen
Liu, Yebin
Computer Vision and Pattern Recognition
This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-realistic quality, and continuous viewpoint control. The process follows a two-stage approach. In the first stage, the system maps audio input to high-fidelity motion sequences, encompassing upper-body gestures and facial expressions. To enrich motion diversity and authenticity, large language model (LLM) priors are integrated with text-aligned semantic audio features, leveraging LLMs' cross-modal generalization power to enhance motion quality. In the second stage, we improve diffusion-based video generation models by incorporating a prior-guided Mixture-of-Experts (MoE) mechanism: a view-guided MoE focuses on view-specific attributes, while a mask-guided MoE enhances region-based rendering stability. Additionally, a mask prediction module is devised to derive human masks from motion data, enhancing the stability and accuracy of masks and enabling mask guiding during inference. We also introduce a comprehensive human video dataset with 2,203 identities, covering diverse body gestures and detailed annotations, facilitating broad generalization. The code, data, and pre-trained models will be released for research purposes.
title Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.23836