DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Peng, Ziqiao, Fan, Yanbo, Wu, Haoyu, Wang, Xuan, Liu, Hongyan, He, Jun, Fan, Zhaoxin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910968489967616
author Peng, Ziqiao
Fan, Yanbo
Wu, Haoyu
Wang, Xuan
Liu, Hongyan
He, Jun
Fan, Zhaoxin
author_facet Peng, Ziqiao
Fan, Yanbo
Wu, Haoyu
Wang, Xuan
Liu, Hongyan
He, Jun
Fan, Zhaoxin
contents In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward transitions. To address this issue, we propose a new task -- multi-round dual-speaker interaction for 3D talking head generation -- which requires models to handle and generate both speaking and listening behaviors in continuous conversation. To solve this task, we introduce DualTalk, a novel unified framework that integrates the dynamic behaviors of speakers and listeners to simulate realistic and coherent dialogue interactions. This framework not only synthesizes lifelike talking heads when speaking but also generates continuous and vivid non-verbal feedback when listening, effectively capturing the interplay between the roles. We also create a new dataset featuring 50 hours of multi-round conversations with over 1,000 characters, where participants continuously switch between speaking and listening roles. Extensive experiments demonstrate that our method significantly enhances the naturalness and expressiveness of 3D talking heads in dual-speaker conversations. We recommend watching the supplementary video: https://ziqiaopeng.github.io/dualtalk.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
Peng, Ziqiao
Fan, Yanbo
Wu, Haoyu
Wang, Xuan
Liu, Hongyan
He, Jun
Fan, Zhaoxin
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward transitions. To address this issue, we propose a new task -- multi-round dual-speaker interaction for 3D talking head generation -- which requires models to handle and generate both speaking and listening behaviors in continuous conversation. To solve this task, we introduce DualTalk, a novel unified framework that integrates the dynamic behaviors of speakers and listeners to simulate realistic and coherent dialogue interactions. This framework not only synthesizes lifelike talking heads when speaking but also generates continuous and vivid non-verbal feedback when listening, effectively capturing the interplay between the roles. We also create a new dataset featuring 50 hours of multi-round conversations with over 1,000 characters, where participants continuously switch between speaking and listening roles. Extensive experiments demonstrate that our method significantly enhances the naturalness and expressiveness of 3D talking heads in dual-speaker conversations. We recommend watching the supplementary video: https://ziqiaopeng.github.io/dualtalk.
title DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
topic Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.18096