The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jia, Wenqi, Liu, Miao, Jiang, Hao, Ananthabhotla, Ishwarya, Rehg, James M., Ithapu, Vamsi Krishna, Gao, Ruohan
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917629242900480
author Jia, Wenqi
Liu, Miao
Jiang, Hao
Ananthabhotla, Ishwarya
Rehg, James M.
Ithapu, Vamsi Krishna
Gao, Ruohan
author_facet Jia, Wenqi
Liu, Miao
Jiang, Hao
Ananthabhotla, Ishwarya
Rehg, James M.
Ithapu, Vamsi Krishna
Gao, Ruohan
contents In recent years, the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions, where both visual and audio signals play a crucial role. While most prior work focus on learning about behaviors that directly involve the camera wearer, we introduce the Ego-Exocentric Conversational Graph Prediction problem, marking the first attempt to infer exocentric conversational interactions from egocentric videos. We propose a unified multi-modal framework -- Audio-Visual Conversational Attention (AV-CONV), for the joint prediction of conversation behaviors -- speaking and listening -- for both the camera wearer as well as all other social partners present in the egocentric video. Specifically, we adopt the self-attention mechanism to model the representations across-time, across-subjects, and across-modalities. To validate our method, we conduct experiments on a challenging egocentric video dataset that includes multi-speaker and multi-conversation scenarios. Our results demonstrate the superior performance of our method compared to a series of baselines. We also present detailed ablation studies to assess the contribution of each component in our model. Check our project page at https://vjwq.github.io/AV-CONV/.
format Preprint
id arxiv_https___arxiv_org_abs_2312_12870
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
Jia, Wenqi
Liu, Miao
Jiang, Hao
Ananthabhotla, Ishwarya
Rehg, James M.
Ithapu, Vamsi Krishna
Gao, Ruohan
Computer Vision and Pattern Recognition
In recent years, the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions, where both visual and audio signals play a crucial role. While most prior work focus on learning about behaviors that directly involve the camera wearer, we introduce the Ego-Exocentric Conversational Graph Prediction problem, marking the first attempt to infer exocentric conversational interactions from egocentric videos. We propose a unified multi-modal framework -- Audio-Visual Conversational Attention (AV-CONV), for the joint prediction of conversation behaviors -- speaking and listening -- for both the camera wearer as well as all other social partners present in the egocentric video. Specifically, we adopt the self-attention mechanism to model the representations across-time, across-subjects, and across-modalities. To validate our method, we conduct experiments on a challenging egocentric video dataset that includes multi-speaker and multi-conversation scenarios. Our results demonstrate the superior performance of our method compared to a series of baselines. We also present detailed ablation studies to assess the contribution of each component in our model. Check our project page at https://vjwq.github.io/AV-CONV/.
title The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.12870