Exploiting temporal information to detect conversational groups in videos and predict the next speaker

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tosato, Lucrezia, Fortier, Victor, Bloch, Isabelle, Pelachaud, Catherine
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914928800038912
author Tosato, Lucrezia
Fortier, Victor
Bloch, Isabelle
Pelachaud, Catherine
author_facet Tosato, Lucrezia
Fortier, Victor
Bloch, Isabelle
Pelachaud, Catherine
contents Studies in human human interaction have introduced the concept of F formation to describe the spatial arrangement of participants during social interactions. This paper has two objectives. It aims at detecting F formations in video sequences and predicting the next speaker in a group conversation. The proposed approach exploits time information and human multimodal signals in video sequences. In particular, we rely on measuring the engagement level of people as a feature of group belonging. Our approach makes use of a recursive neural network, the Long Short Term Memory (LSTM), to predict who will take the speaker's turn in a conversation group. Experiments on the MatchNMingle dataset led to 85% true positives in group detection and 98% accuracy in predicting the next speaker.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16380
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploiting temporal information to detect conversational groups in videos and predict the next speaker
Tosato, Lucrezia
Fortier, Victor
Bloch, Isabelle
Pelachaud, Catherine
Computer Vision and Pattern Recognition
Studies in human human interaction have introduced the concept of F formation to describe the spatial arrangement of participants during social interactions. This paper has two objectives. It aims at detecting F formations in video sequences and predicting the next speaker in a group conversation. The proposed approach exploits time information and human multimodal signals in video sequences. In particular, we rely on measuring the engagement level of people as a feature of group belonging. Our approach makes use of a recursive neural network, the Long Short Term Memory (LSTM), to predict who will take the speaker's turn in a conversation group. Experiments on the MatchNMingle dataset led to 85% true positives in group detection and 98% accuracy in predicting the next speaker.
title Exploiting temporal information to detect conversational groups in videos and predict the next speaker
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.16380