RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Kunyu, Wen, Di, Fu, Jia, Wu, Jiamin, Yang, Kailun, Zheng, Junwei, Liu, Ruiping, Chen, Yufan, Fu, Yuqian, Paudel, Danda Pani, Van Gool, Luc, Stiefelhagen, Rainer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912657662017536
author Peng, Kunyu
Wen, Di
Fu, Jia
Wu, Jiamin
Yang, Kailun
Zheng, Junwei
Liu, Ruiping
Chen, Yufan
Fu, Yuqian
Paudel, Danda Pani
Van Gool, Luc
Stiefelhagen, Rainer
author_facet Peng, Kunyu
Wen, Di
Fu, Jia
Wu, Jiamin
Yang, Kailun
Zheng, Junwei
Liu, Ruiping
Chen, Yufan
Fu, Yuqian
Paudel, Danda Pani
Van Gool, Luc
Stiefelhagen, Rainer
contents Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person of interest conditioned on natural language descriptions. Distinct from conventional action recognition and detection tasks, RAVAR emphasizes precise language-guided action understanding, which is particularly critical for interactive human action analysis in complex multi-person scenarios. In this work, we extend our previously introduced RefAVA dataset to RefAVA++, which comprises >2.9 million frames and >75.1k annotated persons in total. We benchmark this dataset using baselines from multiple related domains, including atomic action localization, video question answering, and text-video retrieval, as well as our earlier model, RefAtomNet. Although RefAtomNet surpasses other baselines by incorporating agent attention to highlight salient features, its ability to align and retrieve cross-modal information remains limited, leading to suboptimal performance in localizing the target person and predicting fine-grained actions. To overcome the aforementioned limitations, we introduce RefAtomNet++, a novel framework that advances cross-modal token aggregation through a multi-hierarchical semantic-aligned cross-attention mechanism combined with multi-trajectory Mamba modeling at the partial-keyword, scene-attribute, and holistic-sentence levels. In particular, scanning trajectories are constructed by dynamically selecting the nearest visual spatial tokens at each timestep for both partial-keyword and scene-attribute levels. Moreover, we design a multi-hierarchical semantic-aligned cross-attention strategy, enabling more effective aggregation of spatial and temporal tokens across different semantic hierarchies. Experiments show that RefAtomNet++ establishes new state-of-the-art results. The dataset and code are released at https://github.com/KPeng9510/refAVA2.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16444
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
Peng, Kunyu
Wen, Di
Fu, Jia
Wu, Jiamin
Yang, Kailun
Zheng, Junwei
Liu, Ruiping
Chen, Yufan
Fu, Yuqian
Paudel, Danda Pani
Van Gool, Luc
Stiefelhagen, Rainer
Computer Vision and Pattern Recognition
Multimedia
Robotics
Image and Video Processing
Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person of interest conditioned on natural language descriptions. Distinct from conventional action recognition and detection tasks, RAVAR emphasizes precise language-guided action understanding, which is particularly critical for interactive human action analysis in complex multi-person scenarios. In this work, we extend our previously introduced RefAVA dataset to RefAVA++, which comprises >2.9 million frames and >75.1k annotated persons in total. We benchmark this dataset using baselines from multiple related domains, including atomic action localization, video question answering, and text-video retrieval, as well as our earlier model, RefAtomNet. Although RefAtomNet surpasses other baselines by incorporating agent attention to highlight salient features, its ability to align and retrieve cross-modal information remains limited, leading to suboptimal performance in localizing the target person and predicting fine-grained actions. To overcome the aforementioned limitations, we introduce RefAtomNet++, a novel framework that advances cross-modal token aggregation through a multi-hierarchical semantic-aligned cross-attention mechanism combined with multi-trajectory Mamba modeling at the partial-keyword, scene-attribute, and holistic-sentence levels. In particular, scanning trajectories are constructed by dynamically selecting the nearest visual spatial tokens at each timestep for both partial-keyword and scene-attribute levels. Moreover, we design a multi-hierarchical semantic-aligned cross-attention strategy, enabling more effective aggregation of spatial and temporal tokens across different semantic hierarchies. Experiments show that RefAtomNet++ establishes new state-of-the-art results. The dataset and code are released at https://github.com/KPeng9510/refAVA2.
title RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
topic Computer Vision and Pattern Recognition
Multimedia
Robotics
Image and Video Processing
url https://arxiv.org/abs/2510.16444