Interaction Region Visual Transformer for Egocentric Action Anticipation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Roy, Debaditya, Rajendiran, Ramanathan, Fernando, Basura
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916087606542336
author Roy, Debaditya
Rajendiran, Ramanathan
Fernando, Basura
author_facet Roy, Debaditya
Rajendiran, Ramanathan
Fernando, Basura
contents Human-object interaction is one of the most important visual cues and we propose a novel way to represent human-object interactions for egocentric action anticipation. We propose a novel transformer variant to model interactions by computing the change in the appearance of objects and human hands due to the execution of the actions and use those changes to refine the video representation. Specifically, we model interactions between hands and objects using Spatial Cross-Attention (SCA) and further infuse contextual information using Trajectory Cross-Attention to obtain environment-refined interaction tokens. Using these tokens, we construct an interaction-centric video representation for action anticipation. We term our model InAViT which achieves state-of-the-art action anticipation performance on large-scale egocentric datasets EPICKTICHENS100 (EK100) and EGTEA Gaze+. InAViT outperforms other visual transformer-based methods including object-centric video representation. On the EK100 evaluation server, InAViT is the top-performing method on the public leaderboard (at the time of submission) where it outperforms the second-best model by 3.3% on mean-top5 recall.
format Preprint
id arxiv_https___arxiv_org_abs_2211_14154
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Interaction Region Visual Transformer for Egocentric Action Anticipation
Roy, Debaditya
Rajendiran, Ramanathan
Fernando, Basura
Computer Vision and Pattern Recognition
Human-object interaction is one of the most important visual cues and we propose a novel way to represent human-object interactions for egocentric action anticipation. We propose a novel transformer variant to model interactions by computing the change in the appearance of objects and human hands due to the execution of the actions and use those changes to refine the video representation. Specifically, we model interactions between hands and objects using Spatial Cross-Attention (SCA) and further infuse contextual information using Trajectory Cross-Attention to obtain environment-refined interaction tokens. Using these tokens, we construct an interaction-centric video representation for action anticipation. We term our model InAViT which achieves state-of-the-art action anticipation performance on large-scale egocentric datasets EPICKTICHENS100 (EK100) and EGTEA Gaze+. InAViT outperforms other visual transformer-based methods including object-centric video representation. On the EK100 evaluation server, InAViT is the top-performing method on the public leaderboard (at the time of submission) where it outperforms the second-best model by 3.3% on mean-top5 recall.
title Interaction Region Visual Transformer for Egocentric Action Anticipation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2211.14154