SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yiheng, Peng, Junran, Shen, Silei, Yang, Jingwei, Wei, ZeJi, Bai, ChenCheng, He, Yonghao, Sui, Wei, Sun, Muyi, Liu, Yan, Yin, Xu-Cheng, Zhang, Man, Zhang, Zhaoxiang, Luo, Chuanchen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915605765947392
author Huang, Yiheng
Peng, Junran
Shen, Silei
Yang, Jingwei
Wei, ZeJi
Bai, ChenCheng
He, Yonghao
Sui, Wei
Sun, Muyi
Liu, Yan
Yin, Xu-Cheng
Zhang, Man
Zhang, Zhaoxiang
Luo, Chuanchen
author_facet Huang, Yiheng
Peng, Junran
Shen, Silei
Yang, Jingwei
Wei, ZeJi
Bai, ChenCheng
He, Yonghao
Sui, Wei
Sun, Muyi
Liu, Yan
Yin, Xu-Cheng
Zhang, Man
Zhang, Zhaoxiang
Luo, Chuanchen
contents The accompanying actions and gestures in dialogue are often closely linked to interactions with the environment, such as looking toward the interlocutor or using gestures to point to the described target at appropriate moments. Speech and semantics guide the production of gestures by determining their timing (WHEN) and style (HOW), while the spatial locations of interactive objects dictate their directional execution (WHERE). Existing approaches either rely solely on descriptive language to generate motions or utilize audio to produce non-interactive gestures, thereby lacking the characterization of interactive timing and spatial intent. This significantly limits the applicability of conversational gesture generation, whether in robotics or in the fields of game and animation production. To address this gap, we present a full-stack solution. We first established a unique data collection method to simultaneously capture high-precision human motion and spatial intent. We then developed a generation model driven by audio, language, and spatial data, alongside dedicated metrics for evaluating interaction timing and spatial accuracy. Finally, we deployed the solution on a humanoid robot, enabling rich, context-aware physical interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
Huang, Yiheng
Peng, Junran
Shen, Silei
Yang, Jingwei
Wei, ZeJi
Bai, ChenCheng
He, Yonghao
Sui, Wei
Sun, Muyi
Liu, Yan
Yin, Xu-Cheng
Zhang, Man
Zhang, Zhaoxiang
Luo, Chuanchen
Graphics
Multimedia
Robotics
The accompanying actions and gestures in dialogue are often closely linked to interactions with the environment, such as looking toward the interlocutor or using gestures to point to the described target at appropriate moments. Speech and semantics guide the production of gestures by determining their timing (WHEN) and style (HOW), while the spatial locations of interactive objects dictate their directional execution (WHERE). Existing approaches either rely solely on descriptive language to generate motions or utilize audio to produce non-interactive gestures, thereby lacking the characterization of interactive timing and spatial intent. This significantly limits the applicability of conversational gesture generation, whether in robotics or in the fields of game and animation production. To address this gap, we present a full-stack solution. We first established a unique data collection method to simultaneously capture high-precision human motion and spatial intent. We then developed a generation model driven by audio, language, and spatial data, alongside dedicated metrics for evaluating interaction timing and spatial accuracy. Finally, we deployed the solution on a humanoid robot, enabling rich, context-aware physical interactions.
title SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
topic Graphics
Multimedia
Robotics
url https://arxiv.org/abs/2509.23852