Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jinhe, Cheng, Yongkang, Hang, Yuming, Han, Gaoge, Li, Jinewei, Zhang, Jing, Gu, Xingjian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909604509646848
author Huang, Jinhe
Cheng, Yongkang
Hang, Yuming
Han, Gaoge
Li, Jinewei
Zhang, Jing
Gu, Xingjian
author_facet Huang, Jinhe
Cheng, Yongkang
Hang, Yuming
Han, Gaoge
Li, Jinewei
Zhang, Jing
Gu, Xingjian
contents Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of listeners in the interaction process and failing to fully explore the dynamic interaction between them. This paper innovatively proposes an Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication. For the first time, we integrate the full-body gestures of listeners into the generation framework. By devising a novel inter-diffusion mechanism, this model can accurately capture the complex interaction patterns between speakers and listeners during communication. In the model construction process, based on the advanced diffusion model architecture, we innovatively introduce interaction conditions and the GAN model to increase the denoising step size. As a result, when generating gesture sequences, the model can not only dynamically generate based on the speaker's speech information but also respond in realtime to the listener's feedback, enabling synergistic interaction between the two. Abundant experimental results demonstrate that compared with the current state-of-the-art gesture generation methods, the model we proposed has achieved remarkable improvements in the naturalness, coherence, and speech-gesture synchronization of the generated gestures. In the subjective evaluation experiments, users highly praised the generated interaction scenarios, believing that they are closer to real life human communication situations. Objective index evaluations also show that our model outperforms the baseline methods in multiple key indicators, providing more powerful support for effective communication.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04996
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication
Huang, Jinhe
Cheng, Yongkang
Hang, Yuming
Han, Gaoge
Li, Jinewei
Zhang, Jing
Gu, Xingjian
Graphics
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of listeners in the interaction process and failing to fully explore the dynamic interaction between them. This paper innovatively proposes an Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication. For the first time, we integrate the full-body gestures of listeners into the generation framework. By devising a novel inter-diffusion mechanism, this model can accurately capture the complex interaction patterns between speakers and listeners during communication. In the model construction process, based on the advanced diffusion model architecture, we innovatively introduce interaction conditions and the GAN model to increase the denoising step size. As a result, when generating gesture sequences, the model can not only dynamically generate based on the speaker's speech information but also respond in realtime to the listener's feedback, enabling synergistic interaction between the two. Abundant experimental results demonstrate that compared with the current state-of-the-art gesture generation methods, the model we proposed has achieved remarkable improvements in the naturalness, coherence, and speech-gesture synchronization of the generated gestures. In the subjective evaluation experiments, users highly praised the generated interaction scenarios, believing that they are closer to real life human communication situations. Objective index evaluations also show that our model outperforms the baseline methods in multiple key indicators, providing more powerful support for effective communication.
title Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication
topic Graphics
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.04996