Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Sangmin, Lai, Bolin, Ryan, Fiona, Boote, Bikram, Rehg, James M.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911856614965248
author Lee, Sangmin
Lai, Bolin
Ryan, Fiona
Boote, Bikram
Rehg, James M.
author_facet Lee, Sangmin
Lai, Bolin
Ryan, Fiona
Boote, Bikram
Rehg, James M.
contents Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-person behaviors or rely on holistic visual representations that are not aligned to utterances in multi-party environments. Consequently, they are limited in modeling the intricate dynamics of multi-party interactions. In this paper, we introduce three new challenging tasks to model the fine-grained dynamics between multiple people: speaking target identification, pronoun coreference resolution, and mentioned player prediction. We contribute extensive data annotations to curate these new challenges in social deduction game settings. Furthermore, we propose a novel multimodal baseline that leverages densely aligned language-visual representations by synchronizing visual features with their corresponding utterances. This facilitates concurrently capturing verbal and non-verbal cues pertinent to social reasoning. Experiments demonstrate the effectiveness of the proposed approach with densely aligned multimodal representations in modeling fine-grained social interactions. Project website: https://sangmin-git.github.io/projects/MMSI.
format Preprint
id arxiv_https___arxiv_org_abs_2403_02090
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
Lee, Sangmin
Lai, Bolin
Ryan, Fiona
Boote, Bikram
Rehg, James M.
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-person behaviors or rely on holistic visual representations that are not aligned to utterances in multi-party environments. Consequently, they are limited in modeling the intricate dynamics of multi-party interactions. In this paper, we introduce three new challenging tasks to model the fine-grained dynamics between multiple people: speaking target identification, pronoun coreference resolution, and mentioned player prediction. We contribute extensive data annotations to curate these new challenges in social deduction game settings. Furthermore, we propose a novel multimodal baseline that leverages densely aligned language-visual representations by synchronizing visual features with their corresponding utterances. This facilitates concurrently capturing verbal and non-verbal cues pertinent to social reasoning. Experiments demonstrate the effectiveness of the proposed approach with densely aligned multimodal representations in modeling fine-grained social interactions. Project website: https://sangmin-git.github.io/projects/MMSI.
title Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2403.02090