CAST: Cross-Attention in Space and Time for Video Action Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lee, Dongho, Lee, Jongseo, Choi, Jinwoo
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917766772031488
author Lee, Dongho
Lee, Jongseo
Choi, Jinwoo
author_facet Lee, Dongho
Lee, Jongseo
Choi, Jinwoo
contents Recognizing human actions in videos requires spatial and temporal understanding. Most existing action recognition models lack a balanced spatio-temporal understanding of videos. In this work, we propose a novel two-stream architecture, called Cross-Attention in Space and Time (CAST), that achieves a balanced spatio-temporal understanding of videos using only RGB input. Our proposed bottleneck cross-attention mechanism enables the spatial and temporal expert models to exchange information and make synergistic predictions, leading to improved performance. We validate the proposed method with extensive experiments on public benchmarks with different characteristics: EPIC-KITCHENS-100, Something-Something-V2, and Kinetics-400. Our method consistently shows favorable performance across these datasets, while the performance of existing methods fluctuates depending on the dataset characteristics.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18825
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CAST: Cross-Attention in Space and Time for Video Action Recognition
Lee, Dongho
Lee, Jongseo
Choi, Jinwoo
Computer Vision and Pattern Recognition
Recognizing human actions in videos requires spatial and temporal understanding. Most existing action recognition models lack a balanced spatio-temporal understanding of videos. In this work, we propose a novel two-stream architecture, called Cross-Attention in Space and Time (CAST), that achieves a balanced spatio-temporal understanding of videos using only RGB input. Our proposed bottleneck cross-attention mechanism enables the spatial and temporal expert models to exchange information and make synergistic predictions, leading to improved performance. We validate the proposed method with extensive experiments on public benchmarks with different characteristics: EPIC-KITCHENS-100, Something-Something-V2, and Kinetics-400. Our method consistently shows favorable performance across these datasets, while the performance of existing methods fluctuates depending on the dataset characteristics.
title CAST: Cross-Attention in Space and Time for Video Action Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.18825