Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Yuhang, Tang, Zixuan, Pang, Yunsheng, Ding, Beichen, Liu, Mengyuan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917560906153984
author Wen, Yuhang
Tang, Zixuan
Pang, Yunsheng
Ding, Beichen
Liu, Mengyuan
author_facet Wen, Yuhang
Tang, Zixuan
Pang, Yunsheng
Ding, Beichen
Liu, Mengyuan
contents Recognizing interactive action plays an important role in human-robot interaction and collaboration. Previous methods use late fusion and co-attention mechanism to capture interactive relations, which have limited learning capability or inefficiency to adapt to more interacting entities. With assumption that priors of each entity are already known, they also lack evaluations on a more general setting addressing the diversity of subjects. To address these problems, we propose an Interactive Spatiotemporal Token Attention Network (ISTA-Net), which simultaneously model spatial, temporal, and interactive relations. Specifically, our network contains a tokenizer to partition Interactive Spatiotemporal Tokens (ISTs), which is a unified way to represent motions of multiple diverse entities. By extending the entity dimension, ISTs provide better interactive representations. To jointly learn along three dimensions in ISTs, multi-head self-attention blocks integrated with 3D convolutions are designed to capture inter-token correlations. When modeling correlations, a strict entity ordering is usually irrelevant for recognizing interactive actions. To this end, Entity Rearrangement is proposed to eliminate the orderliness in ISTs for interchangeable entities. Extensive experiments on four datasets verify the effectiveness of ISTA-Net by outperforming state-of-the-art methods. Our code is publicly available at https://github.com/Necolizer/ISTA-Net
format Preprint
id arxiv_https___arxiv_org_abs_2307_07469
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
Wen, Yuhang
Tang, Zixuan
Pang, Yunsheng
Ding, Beichen
Liu, Mengyuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Recognizing interactive action plays an important role in human-robot interaction and collaboration. Previous methods use late fusion and co-attention mechanism to capture interactive relations, which have limited learning capability or inefficiency to adapt to more interacting entities. With assumption that priors of each entity are already known, they also lack evaluations on a more general setting addressing the diversity of subjects. To address these problems, we propose an Interactive Spatiotemporal Token Attention Network (ISTA-Net), which simultaneously model spatial, temporal, and interactive relations. Specifically, our network contains a tokenizer to partition Interactive Spatiotemporal Tokens (ISTs), which is a unified way to represent motions of multiple diverse entities. By extending the entity dimension, ISTs provide better interactive representations. To jointly learn along three dimensions in ISTs, multi-head self-attention blocks integrated with 3D convolutions are designed to capture inter-token correlations. When modeling correlations, a strict entity ordering is usually irrelevant for recognizing interactive actions. To this end, Entity Rearrangement is proposed to eliminate the orderliness in ISTs for interchangeable entities. Extensive experiments on four datasets verify the effectiveness of ISTA-Net by outperforming state-of-the-art methods. Our code is publicly available at https://github.com/Necolizer/ISTA-Net
title Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2307.07469