Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies
Fuente:
arXiv
Saved in:
| Main Authors: | Gorostegui, Juan Ignacio Bustos, Buemi, Maria Elena |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Body-Hand Modality Expertized Networks with Cross-attention for Fine-grained Skeleton Action Recognition
by: Cho, Seungyeon, et al.
Published: (2025)
by: Cho, Seungyeon, et al.
Published: (2025)
Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
by: Tse, Tze Ho Elden, et al.
Published: (2025)
by: Tse, Tze Ho Elden, et al.
Published: (2025)
Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
by: Liu, Chao, et al.
Published: (2024)
by: Liu, Chao, et al.
Published: (2024)
Real-Time Hand Gesture Recognition: Integrating Skeleton-Based Data Fusion and Multi-Stream CNN
by: Yusuf, Oluwaleke, et al.
Published: (2024)
by: Yusuf, Oluwaleke, et al.
Published: (2024)
ARN-LSTM: A Multi-Stream Fusion Model for Skeleton-based Action Recognition
by: Wang, Chuanchuan, et al.
Published: (2024)
by: Wang, Chuanchuan, et al.
Published: (2024)
MambaSOD: Dual Mamba-Driven Cross-Modal Fusion Network for RGB-D Salient Object Detection
by: Zhan, Yue, et al.
Published: (2024)
by: Zhan, Yue, et al.
Published: (2024)
Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion
by: Hyder, Syed Waleed, et al.
Published: (2023)
by: Hyder, Syed Waleed, et al.
Published: (2023)
BHaRNet: Reliability-Aware Body-Hand Modality Expertized Networks for Fine-grained Skeleton Action Recognition
by: Cho, Seungyeon, et al.
Published: (2026)
by: Cho, Seungyeon, et al.
Published: (2026)
Local Spherical Harmonics Improve Skeleton-Based Hand Action Recognition
by: Prasse, Katharina, et al.
Published: (2023)
by: Prasse, Katharina, et al.
Published: (2023)
SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition
by: Huang, Shunyu, et al.
Published: (2026)
by: Huang, Shunyu, et al.
Published: (2026)
In My Perspective, In My Hands: Accurate Egocentric 2D Hand Pose and Action Recognition
by: Mucha, Wiktor, et al.
Published: (2024)
by: Mucha, Wiktor, et al.
Published: (2024)
MADiff: Motion-Aware Mamba Diffusion Models for Hand Trajectory Prediction on Egocentric Videos
by: Ma, Junyi, et al.
Published: (2024)
by: Ma, Junyi, et al.
Published: (2024)
Recovering Complete Actions for Cross-dataset Skeleton Action Recognition
by: Liu, Hanchao, et al.
Published: (2024)
by: Liu, Hanchao, et al.
Published: (2024)
Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification
by: Nguyen, Y Hop, et al.
Published: (2025)
by: Nguyen, Y Hop, et al.
Published: (2025)
Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities
by: Santos-Villafranca, Maria, et al.
Published: (2025)
by: Santos-Villafranca, Maria, et al.
Published: (2025)
Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
Expressive Keypoints for Skeleton-based Action Recognition via Skeleton Transformation
by: Yang, Yijie, et al.
Published: (2024)
by: Yang, Yijie, et al.
Published: (2024)
SkeletonX: Data-Efficient Skeleton-based Action Recognition via Cross-sample Feature Aggregation
by: Zhang, Zongye, et al.
Published: (2025)
by: Zhang, Zongye, et al.
Published: (2025)
TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
by: Liu, Yanan, et al.
Published: (2025)
by: Liu, Yanan, et al.
Published: (2025)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
by: Kukleva, Anna, et al.
Published: (2024)
by: Kukleva, Anna, et al.
Published: (2024)
A Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective
by: Truong, Thanh-Dat, et al.
Published: (2023)
by: Truong, Thanh-Dat, et al.
Published: (2023)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
by: Zhang, Mingfang, et al.
Published: (2024)
by: Zhang, Mingfang, et al.
Published: (2024)
Heatmap Pooling Network for Action Recognition from RGB Videos
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
HFGCN:Hypergraph Fusion Graph Convolutional Networks for Skeleton-Based Action Recognition
by: Dong, Pengcheng, et al.
Published: (2025)
by: Dong, Pengcheng, et al.
Published: (2025)
SkelMamba: A State Space Model for Efficient Skeleton Action Recognition of Neurological Disorders
by: Martinel, Niki, et al.
Published: (2024)
by: Martinel, Niki, et al.
Published: (2024)
SHARP: Segmentation of Hands and Arms by Range using Pseudo-Depth for Enhanced Egocentric 3D Hand Pose Estimation and Action Recognition
by: Mucha, Wiktor, et al.
Published: (2024)
by: Mucha, Wiktor, et al.
Published: (2024)
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
by: Marouani, Alexis, et al.
Published: (2026)
by: Marouani, Alexis, et al.
Published: (2026)
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
by: Wen, Yuhang, et al.
Published: (2023)
by: Wen, Yuhang, et al.
Published: (2023)
SBF: An Effective Representation to Augment Skeleton for Video-based Human Action Recognition
by: Peng, Zhuoxuan, et al.
Published: (2026)
by: Peng, Zhuoxuan, et al.
Published: (2026)
Adaptive Physical-Facial Representation Fusion via Subject-Invariant Cross-Modal Prompt Tuning for Video-Based Emotion Recognition
by: Luo, Xiwen, et al.
Published: (2026)
by: Luo, Xiwen, et al.
Published: (2026)
SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
by: Liu, Hongda, et al.
Published: (2025)
by: Liu, Hongda, et al.
Published: (2025)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
by: Wei, Yiping, et al.
Published: (2023)
by: Wei, Yiping, et al.
Published: (2023)
Detecting Precise Hand Touch Moments in Egocentric Video
by: Nguyen, Huy Anh, et al.
Published: (2026)
by: Nguyen, Huy Anh, et al.
Published: (2026)
Complementing Event Streams and RGB Frames for Hand Mesh Reconstruction
by: Jiang, Jianping, et al.
Published: (2024)
by: Jiang, Jianping, et al.
Published: (2024)
Sign Language Recognition Based On Facial Expression and Hand Skeleton
by: Long, Zhiyu, et al.
Published: (2024)
by: Long, Zhiyu, et al.
Published: (2024)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
by: Zhu, Haodong, et al.
Published: (2025)
by: Zhu, Haodong, et al.
Published: (2025)
Similar Items
-
Body-Hand Modality Expertized Networks with Cross-attention for Fine-grained Skeleton Action Recognition
by: Cho, Seungyeon, et al.
Published: (2025) -
Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
by: Tse, Tze Ho Elden, et al.
Published: (2025) -
Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
by: Liu, Chao, et al.
Published: (2024) -
Real-Time Hand Gesture Recognition: Integrating Skeleton-Based Data Fusion and Multi-Stream CNN
by: Yusuf, Oluwaleke, et al.
Published: (2024) -
ARN-LSTM: A Multi-Stream Fusion Model for Skeleton-based Action Recognition
by: Wang, Chuanchuan, et al.
Published: (2024)