Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deichler, Anna, Wang, Siyang, Alexanderson, Simon, Beskow, Jonas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912588769525760
author Deichler, Anna
Wang, Siyang
Alexanderson, Simon
Beskow, Jonas
author_facet Deichler, Anna
Wang, Siyang
Alexanderson, Simon
Beskow, Jonas
contents One of the main goals of robotics and intelligent agent research is to enable natural communication with humans in physically situated settings. While recent work has focused on verbal modes such as language and speech, non-verbal communication is crucial for flexible interaction. We present a framework for generating pointing gestures in embodied agents by combining imitation and reinforcement learning. Using a small motion capture dataset, our method learns a motor control policy that produces physically valid, naturalistic gestures with high referential accuracy. We evaluate the approach against supervised learning and retrieval baselines in both objective metrics and a virtual reality referential game with human users. Results show that our system achieves higher naturalness and accuracy than state-of-the-art supervised models, highlighting the promise of imitation-RL for communicative gesture generation and its potential application to robots.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
Deichler, Anna
Wang, Siyang
Alexanderson, Simon
Beskow, Jonas
Robotics
Human-Computer Interaction
Machine Learning
68T07, 68T40
I.2.9; I.2.6
One of the main goals of robotics and intelligent agent research is to enable natural communication with humans in physically situated settings. While recent work has focused on verbal modes such as language and speech, non-verbal communication is crucial for flexible interaction. We present a framework for generating pointing gestures in embodied agents by combining imitation and reinforcement learning. Using a small motion capture dataset, our method learns a motor control policy that produces physically valid, naturalistic gestures with high referential accuracy. We evaluate the approach against supervised learning and retrieval baselines in both objective metrics and a virtual reality referential game with human users. Results show that our system achieves higher naturalness and accuracy than state-of-the-art supervised models, highlighting the promise of imitation-RL for communicative gesture generation and its potential application to robots.
title Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
topic Robotics
Human-Computer Interaction
Machine Learning
68T07, 68T40
I.2.9; I.2.6
url https://arxiv.org/abs/2509.12507