Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lai, Bolin, Ryan, Fiona, Jia, Wenqi, Liu, Miao, Rehg, James M.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917619604389888
author Lai, Bolin
Ryan, Fiona
Jia, Wenqi
Liu, Miao
Rehg, James M.
author_facet Lai, Bolin
Ryan, Fiona
Jia, Wenqi
Liu, Miao
Rehg, James M.
contents Egocentric gaze anticipation serves as a key building block for the emerging capability of Augmented Reality. Notably, gaze behavior is driven by both visual cues and audio signals during daily activities. Motivated by this observation, we introduce the first model that leverages both the video and audio modalities for egocentric gaze anticipation. Specifically, we propose a Contrastive Spatial-Temporal Separable (CSTS) fusion approach that adopts two modules to separately capture audio-visual correlations in spatial and temporal dimensions, and applies a contrastive loss on the re-weighted audio-visual features from fusion modules for representation learning. We conduct extensive ablation studies and thorough analysis using two egocentric video datasets: Ego4D and Aria, to validate our model design. We demonstrate the audio improves the performance by +2.5% and +2.4% on the two datasets. Our model also outperforms the prior state-of-the-art methods by at least +1.9% and +1.6%. Moreover, we provide visualizations to show the gaze anticipation results and provide additional insights into audio-visual representation learning. The code and data split are available on our website (https://bolinlai.github.io/CSTS-EgoGazeAnticipation/).
format Preprint
id arxiv_https___arxiv_org_abs_2305_03907
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
Lai, Bolin
Ryan, Fiona
Jia, Wenqi
Liu, Miao
Rehg, James M.
Computer Vision and Pattern Recognition
Egocentric gaze anticipation serves as a key building block for the emerging capability of Augmented Reality. Notably, gaze behavior is driven by both visual cues and audio signals during daily activities. Motivated by this observation, we introduce the first model that leverages both the video and audio modalities for egocentric gaze anticipation. Specifically, we propose a Contrastive Spatial-Temporal Separable (CSTS) fusion approach that adopts two modules to separately capture audio-visual correlations in spatial and temporal dimensions, and applies a contrastive loss on the re-weighted audio-visual features from fusion modules for representation learning. We conduct extensive ablation studies and thorough analysis using two egocentric video datasets: Ego4D and Aria, to validate our model design. We demonstrate the audio improves the performance by +2.5% and +2.4% on the two datasets. Our model also outperforms the prior state-of-the-art methods by at least +1.9% and +1.6%. Moreover, we provide visualizations to show the gaze anticipation results and provide additional insights into audio-visual representation learning. The code and data split are available on our website (https://bolinlai.github.io/CSTS-EgoGazeAnticipation/).
title Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.03907