Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Suresh, Varsha, Abootorabi, Mohammad Mahdi, Salman, Mohamed, Mughal, M. Hamza, Theobalt, Christian, Ram, Ashwin, Steimle, Jürgen, Demberg, Vera
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916073740173312
author Suresh, Varsha
Abootorabi, Mohammad Mahdi
Salman, Mohamed
Mughal, M. Hamza
Theobalt, Christian
Ram, Ashwin
Steimle, Jürgen
Demberg, Vera
author_facet Suresh, Varsha
Abootorabi, Mohammad Mahdi
Salman, Mohamed
Mughal, M. Hamza
Theobalt, Christian
Ram, Ashwin
Steimle, Jürgen
Demberg, Vera
contents Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically meaningful gestures whose communicative intent is not captured by motion alone. Direct contrastive alignment between transcripts and continuous motion embeddings often overemphasizes low-level kinematics and misses the symbolic content of semantic gestures. We propose semantic motion anchors, natural-language abstractions of gesture motion capturing physical form and communicative intent. Our method discretizes 3D gestures into body-hand motion primitives, verbalizes them into structured descriptions, and grounds them in the transcript to provide auxiliary contrastive supervision. On BEAT2, our method improves text-to-gesture R@1 by 8.2% over a direct text-motion baseline and outperforms prior retrieval approaches on text to gesture and gesture to text retrieval directions. Beyond aggregate retrieval metrics, semantic motion anchor supervision helps retrieve gestures that are semantically meaningful for the spoken query, rather than defaulting to generic motion patterns. A downstream retrieval-augmented gesture generation study showed that users significantly preferred gestures retrieved by our approach over a retrieval-augmented generation baseline, demonstrating that semantically grounded retrieval translates to gestures that better convey communicative intent in downstream generation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30608
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
Suresh, Varsha
Abootorabi, Mohammad Mahdi
Salman, Mohamed
Mughal, M. Hamza
Theobalt, Christian
Ram, Ashwin
Steimle, Jürgen
Demberg, Vera
Computation and Language
Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically meaningful gestures whose communicative intent is not captured by motion alone. Direct contrastive alignment between transcripts and continuous motion embeddings often overemphasizes low-level kinematics and misses the symbolic content of semantic gestures. We propose semantic motion anchors, natural-language abstractions of gesture motion capturing physical form and communicative intent. Our method discretizes 3D gestures into body-hand motion primitives, verbalizes them into structured descriptions, and grounds them in the transcript to provide auxiliary contrastive supervision. On BEAT2, our method improves text-to-gesture R@1 by 8.2% over a direct text-motion baseline and outperforms prior retrieval approaches on text to gesture and gesture to text retrieval directions. Beyond aggregate retrieval metrics, semantic motion anchor supervision helps retrieve gestures that are semantically meaningful for the spoken query, rather than defaulting to generic motion patterns. A downstream retrieval-augmented gesture generation study showed that users significantly preferred gestures retrieved by our approach over a retrieval-augmented generation baseline, demonstrating that semantically grounded retrieval translates to gestures that better convey communicative intent in downstream generation.
title Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
topic Computation and Language
url https://arxiv.org/abs/2605.30608