Exploiting temporal information to detect conversational groups in videos and predict the next speaker
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914928800038912 |
|---|---|
| author | Tosato, Lucrezia Fortier, Victor Bloch, Isabelle Pelachaud, Catherine |
| author_facet | Tosato, Lucrezia Fortier, Victor Bloch, Isabelle Pelachaud, Catherine |
| contents | Studies in human human interaction have introduced the concept of F formation to describe the spatial arrangement of participants during social interactions. This paper has two objectives. It aims at detecting F formations in video sequences and predicting the next speaker in a group conversation. The proposed approach exploits time information and human multimodal signals in video sequences. In particular, we rely on measuring the engagement level of people as a feature of group belonging. Our approach makes use of a recursive neural network, the Long Short Term Memory (LSTM), to predict who will take the speaker's turn in a conversation group. Experiments on the MatchNMingle dataset led to 85% true positives in group detection and 98% accuracy in predicting the next speaker. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_16380 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Exploiting temporal information to detect conversational groups in videos and predict the next speaker Tosato, Lucrezia Fortier, Victor Bloch, Isabelle Pelachaud, Catherine Computer Vision and Pattern Recognition Studies in human human interaction have introduced the concept of F formation to describe the spatial arrangement of participants during social interactions. This paper has two objectives. It aims at detecting F formations in video sequences and predicting the next speaker in a group conversation. The proposed approach exploits time information and human multimodal signals in video sequences. In particular, we rely on measuring the engagement level of people as a feature of group belonging. Our approach makes use of a recursive neural network, the Long Short Term Memory (LSTM), to predict who will take the speaker's turn in a conversation group. Experiments on the MatchNMingle dataset led to 85% true positives in group detection and 98% accuracy in predicting the next speaker. |
| title | Exploiting temporal information to detect conversational groups in videos and predict the next speaker |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2408.16380 |