Recognizing Co-Speech Gestures in-the-Wild

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hegde, Sindhu B, Prajwal, K R, Zisserman, Andrew
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914618299908096
author Hegde, Sindhu B
Prajwal, K R
Zisserman, Andrew
author_facet Hegde, Sindhu B
Prajwal, K R
Zisserman, Andrew
contents While humans naturally gesture during speech, only a sparse subset of these movements are visually depictive and semantically linked to specific spoken words. Current multimodal models struggle to capture these semantic co-speech gestures, heavily bottlenecked by a lack of precisely annotated training data. To address this, we introduce the Gesture Recognition in the Wild (GRW) dataset, the first large-scale benchmark designed to map unconstrained human gestures to specific words with frame-accurate temporal boundaries. Comprising 156,688 manually annotated video clips, GRW spans a highly diverse 150-word taxonomy of physical actions, spatial descriptors, and abstract concepts. We leverage GRW to train video models to (a) classify gestures as semantic or not, (b) recognize the word corresponding to a co-speech gesture, and (c) temporally localize the gesture. We also use GRW to establish benchmarks for these three tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31589
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Recognizing Co-Speech Gestures in-the-Wild
Hegde, Sindhu B
Prajwal, K R
Zisserman, Andrew
Computer Vision and Pattern Recognition
While humans naturally gesture during speech, only a sparse subset of these movements are visually depictive and semantically linked to specific spoken words. Current multimodal models struggle to capture these semantic co-speech gestures, heavily bottlenecked by a lack of precisely annotated training data. To address this, we introduce the Gesture Recognition in the Wild (GRW) dataset, the first large-scale benchmark designed to map unconstrained human gestures to specific words with frame-accurate temporal boundaries. Comprising 156,688 manually annotated video clips, GRW spans a highly diverse 150-word taxonomy of physical actions, spatial descriptors, and abstract concepts. We leverage GRW to train video models to (a) classify gestures as semantic or not, (b) recognize the word corresponding to a co-speech gesture, and (c) temporally localize the gesture. We also use GRW to establish benchmarks for these three tasks.
title Recognizing Co-Speech Gestures in-the-Wild
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.31589