Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Lokesh, Shah, Nirmesh, Gudmalwar, Ashishkumar P., Wasnik, Pankaj
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910061249429504
author Kumar, Lokesh
Shah, Nirmesh
Gudmalwar, Ashishkumar P.
Wasnik, Pankaj
author_facet Kumar, Lokesh
Shah, Nirmesh
Gudmalwar, Ashishkumar P.
Wasnik, Pankaj
contents Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent text-to-speech (TTS) systems have begun incorporating multimodal cues such as facial expressions or lip movements, the role of hand gestures in shaping prosody remains largely underexplored. We propose a novel multimodal TTS framework, Gesture2Speech, that leverages visual gesture cues to modulate prosody in synthesized speech. Motivated by the observation that confident and expressive speakers coordinate gestures with vocal prosody, we introduce a multimodal Mixture-of-Experts (MoE) architecture that dynamically fuses linguistic content and gesture features within a dedicated style extraction module. The fused representation conditions an LLM-based speech decoder, enabling prosodic modulation that is temporally aligned with hand movements. We further design a gesture-speech alignment loss that explicitly models their temporal correspondence to ensure fine-grained synchrony between gestures and prosodic contours. Evaluations on the PATS dataset show that Gesture2Speech outperforms state-of-the-art baselines in both speech naturalness and gesture-speech synchrony. To the best of our knowledge, this is the first work to utilize hand gesture cues for prosody control in neural speech synthesis. Demo samples are available at https://research.sri-media-analysis.com/aaai26-beeu-gesture2speech/
format Preprint
id arxiv_https___arxiv_org_abs_2603_19831
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
Kumar, Lokesh
Shah, Nirmesh
Gudmalwar, Ashishkumar P.
Wasnik, Pankaj
Audio and Speech Processing
Artificial Intelligence
Multimedia
Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent text-to-speech (TTS) systems have begun incorporating multimodal cues such as facial expressions or lip movements, the role of hand gestures in shaping prosody remains largely underexplored. We propose a novel multimodal TTS framework, Gesture2Speech, that leverages visual gesture cues to modulate prosody in synthesized speech. Motivated by the observation that confident and expressive speakers coordinate gestures with vocal prosody, we introduce a multimodal Mixture-of-Experts (MoE) architecture that dynamically fuses linguistic content and gesture features within a dedicated style extraction module. The fused representation conditions an LLM-based speech decoder, enabling prosodic modulation that is temporally aligned with hand movements. We further design a gesture-speech alignment loss that explicitly models their temporal correspondence to ensure fine-grained synchrony between gestures and prosodic contours. Evaluations on the PATS dataset show that Gesture2Speech outperforms state-of-the-art baselines in both speech naturalness and gesture-speech synchrony. To the best of our knowledge, this is the first work to utilize hand gesture cues for prosody control in neural speech synthesis. Demo samples are available at https://research.sri-media-analysis.com/aaai26-beeu-gesture2speech/
title Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
topic Audio and Speech Processing
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2603.19831