MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yamane, Taiga, Suzuki, Satoshi, Masumura, Ryo, Tora, Shotaro
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914135572217856
author Yamane, Taiga
Suzuki, Satoshi
Masumura, Ryo
Tora, Shotaro
author_facet Yamane, Taiga
Suzuki, Satoshi
Masumura, Ryo
Tora, Shotaro
contents Multi-view action recognition aims to recognize human actions using multiple camera views and deals with occlusion caused by obstacles or crowds. In this task, cooperation among views, which generates a joint representation by combining multiple views, is vital. Previous studies have explored promising cooperation methods for improving performance. However, since their methods focus only on the task setting of recognizing a single action from an entire video, they are not applicable to the recently popular spatio-temporal action recognition~(STAR) setting, in which each person's action is recognized sequentially. To address this problem, this paper proposes a multi-view action recognition method for the STAR setting, called MVAFormer. In MVAFormer, we introduce a novel transformer-based cooperation module among views. In contrast to previous studies, which utilize embedding vectors with lost spatial information, our module utilizes the feature map for effective cooperation in the STAR setting, which preserves the spatial information. Furthermore, in our module, we divide the self-attention for the same and different views to model the relationship between multiple views effectively. The results of experiments using a newly collected dataset demonstrate that MVAFormer outperforms the comparison baselines by approximately $4.4$ points on the F-measure.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
Yamane, Taiga
Suzuki, Satoshi
Masumura, Ryo
Tora, Shotaro
Computer Vision and Pattern Recognition
Multi-view action recognition aims to recognize human actions using multiple camera views and deals with occlusion caused by obstacles or crowds. In this task, cooperation among views, which generates a joint representation by combining multiple views, is vital. Previous studies have explored promising cooperation methods for improving performance. However, since their methods focus only on the task setting of recognizing a single action from an entire video, they are not applicable to the recently popular spatio-temporal action recognition~(STAR) setting, in which each person's action is recognized sequentially. To address this problem, this paper proposes a multi-view action recognition method for the STAR setting, called MVAFormer. In MVAFormer, we introduce a novel transformer-based cooperation module among views. In contrast to previous studies, which utilize embedding vectors with lost spatial information, our module utilizes the feature map for effective cooperation in the STAR setting, which preserves the spatial information. Furthermore, in our module, we divide the self-attention for the same and different views to model the relationship between multiple views effectively. The results of experiments using a newly collected dataset demonstrate that MVAFormer outperforms the comparison baselines by approximately $4.4$ points on the F-measure.
title MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.02473