Co-Speech Gesture Detection through Multi-Phase Sequence Labeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghaleb, Esam, Burenko, Ilya, Rasenberg, Marlou, Pouw, Wim, Uhrig, Peter, Holler, Judith, Toni, Ivan, Özyürek, Aslı, Fernández, Raquel
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916225612775424
author Ghaleb, Esam
Burenko, Ilya
Rasenberg, Marlou
Pouw, Wim
Uhrig, Peter
Holler, Judith
Toni, Ivan
Özyürek, Aslı
Fernández, Raquel
author_facet Ghaleb, Esam
Burenko, Ilya
Rasenberg, Marlou
Pouw, Wim
Uhrig, Peter
Holler, Judith
Toni, Ivan
Özyürek, Aslı
Fernández, Raquel
contents Gestures are integral components of face-to-face communication. They unfold over time, often following predictable movement phases of preparation, stroke, and retraction. Yet, the prevalent approach to automatic gesture detection treats the problem as binary classification, classifying a segment as either containing a gesture or not, thus failing to capture its inherently sequential and contextual nature. To address this, we introduce a novel framework that reframes the task as a multi-phase sequence labeling problem rather than binary classification. Our model processes sequences of skeletal movements over time windows, uses Transformer encoders to learn contextual embeddings, and leverages Conditional Random Fields to perform sequence labeling. We evaluate our proposal on a large dataset of diverse co-speech gestures in task-oriented face-to-face dialogues. The results consistently demonstrate that our method significantly outperforms strong baseline models in detecting gesture strokes. Furthermore, applying Transformer encoders to learn contextual embeddings from movement sequences substantially improves gesture unit detection. These results highlight our framework's capacity to capture the fine-grained dynamics of co-speech gesture phases, paving the way for more nuanced and accurate gesture detection and analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2308_10680
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Co-Speech Gesture Detection through Multi-Phase Sequence Labeling
Ghaleb, Esam
Burenko, Ilya
Rasenberg, Marlou
Pouw, Wim
Uhrig, Peter
Holler, Judith
Toni, Ivan
Özyürek, Aslı
Fernández, Raquel
Computer Vision and Pattern Recognition
Gestures are integral components of face-to-face communication. They unfold over time, often following predictable movement phases of preparation, stroke, and retraction. Yet, the prevalent approach to automatic gesture detection treats the problem as binary classification, classifying a segment as either containing a gesture or not, thus failing to capture its inherently sequential and contextual nature. To address this, we introduce a novel framework that reframes the task as a multi-phase sequence labeling problem rather than binary classification. Our model processes sequences of skeletal movements over time windows, uses Transformer encoders to learn contextual embeddings, and leverages Conditional Random Fields to perform sequence labeling. We evaluate our proposal on a large dataset of diverse co-speech gestures in task-oriented face-to-face dialogues. The results consistently demonstrate that our method significantly outperforms strong baseline models in detecting gesture strokes. Furthermore, applying Transformer encoders to learn contextual embeddings from movement sequences substantially improves gesture unit detection. These results highlight our framework's capacity to capture the fine-grained dynamics of co-speech gesture phases, paving the way for more nuanced and accurate gesture detection and analysis.
title Co-Speech Gesture Detection through Multi-Phase Sequence Labeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.10680