A Survey on Backbones for Deep Video Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Zixuan, Zhao, Youjun, Wen, Yuhang, Liu, Mengyuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
by: Wen, Yuhang, et al.
Published: (2023)
by: Wen, Yuhang, et al.
Published: (2023)
HDBN: A Novel Hybrid Dual-branch Network for Robust Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition
by: Liu, Mengyuan, et al.
Published: (2026)
by: Liu, Mengyuan, et al.
Published: (2026)
Exploring Explainability in Video Action Recognition
by: Saha, Avinab, et al.
Published: (2024)
by: Saha, Avinab, et al.
Published: (2024)
VG4D: Vision-Language Model Goes 4D Video Recognition
by: Deng, Zhichao, et al.
Published: (2024)
by: Deng, Zhichao, et al.
Published: (2024)
Flatten: Video Action Recognition is an Image Classification task
by: Chen, Junlin, et al.
Published: (2024)
by: Chen, Junlin, et al.
Published: (2024)
SkateboardAI: The Coolest Video Action Recognition for Skateboarding
by: Chen, Hanxiao
Published: (2023)
by: Chen, Hanxiao
Published: (2023)
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025)
by: Kim, Joochan, et al.
Published: (2025)
SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation
by: Tang, Zixuan, et al.
Published: (2026)
by: Tang, Zixuan, et al.
Published: (2026)
Deep Learning in Palmprint Recognition-A Comprehensive Survey
by: Gao, Chengrui, et al.
Published: (2025)
by: Gao, Chengrui, et al.
Published: (2025)
Align before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action Recognition
by: Chen, Yifei, et al.
Published: (2023)
by: Chen, Yifei, et al.
Published: (2023)
Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition
by: Ni, Xinzhe, et al.
Published: (2022)
by: Ni, Xinzhe, et al.
Published: (2022)
Explore Human Parsing Modality for Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
SV3.3B: A Sports Video Understanding Model for Action Recognition
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition
by: Liu, Hongli, et al.
Published: (2026)
by: Liu, Hongli, et al.
Published: (2026)
Fire on Motion: Optimizing Video Pass-bands for Efficient Spiking Action Recognition
by: Ye, Shuhan, et al.
Published: (2026)
by: Ye, Shuhan, et al.
Published: (2026)
Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
by: Mago, Gowreesh, et al.
Published: (2025)
by: Mago, Gowreesh, et al.
Published: (2025)
SAM2 for Image and Video Segmentation: A Comprehensive Survey
by: Jiaxing, Zhang, et al.
Published: (2025)
by: Jiaxing, Zhang, et al.
Published: (2025)
Improving Skeleton-based Action Recognition with Interactive Object Information
by: Wen, Hao, et al.
Published: (2025)
by: Wen, Hao, et al.
Published: (2025)
Interpretable Action Recognition on Hard to Classify Actions
by: Anichenko, Anastasia, et al.
Published: (2024)
by: Anichenko, Anastasia, et al.
Published: (2024)
EPAM-Net: An Efficient Pose-driven Attention-guided Multimodal Network for Video Action Recognition
by: Abdelkawy, Ahmed, et al.
Published: (2024)
by: Abdelkawy, Ahmed, et al.
Published: (2024)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
by: Kang, Seokun, et al.
Published: (2025)
by: Kang, Seokun, et al.
Published: (2025)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022)
by: Li, Bing, et al.
Published: (2022)
Vamos: Versatile Action Models for Video Understanding
by: Wang, Shijie, et al.
Published: (2023)
by: Wang, Shijie, et al.
Published: (2023)
Efficient Spatial-Temporal Modeling for Real-Time Video Analysis: A Unified Framework for Action Recognition and Object Tracking
by: John, Shahla
Published: (2025)
by: John, Shahla
Published: (2025)
TASAR: Transfer-based Attack on Skeletal Action Recognition
by: Diao, Yunfeng, et al.
Published: (2024)
by: Diao, Yunfeng, et al.
Published: (2024)
Idempotent Unsupervised Representation Learning for Skeleton-Based Action Recognition
by: Lin, Lilang, et al.
Published: (2024)
by: Lin, Lilang, et al.
Published: (2024)
Handwritten Text Recognition: A Survey
by: Garrido-Munoz, Carlos, et al.
Published: (2025)
by: Garrido-Munoz, Carlos, et al.
Published: (2025)
A Survey on Video Anomaly Detection via Deep Learning: Human, Vehicle, and Environment
by: Noghre, Ghazal Alinezhad, et al.
Published: (2025)
by: Noghre, Ghazal Alinezhad, et al.
Published: (2025)
Towards A Comprehensive Visual Saliency Explanation Framework for AI-based Face Recognition Systems
by: Lu, Yuhang, et al.
Published: (2024)
by: Lu, Yuhang, et al.
Published: (2024)
CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition
by: Wen, Yuhang, et al.
Published: (2024)
by: Wen, Yuhang, et al.
Published: (2024)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
Advancing Complex Video Object Segmentation via Progressive Concept Construction
by: Zhang, Zhixiong, et al.
Published: (2025)
by: Zhang, Zhixiong, et al.
Published: (2025)
Revisiting the Integration of Convolution and Attention for Vision Backbone
by: Zhu, Lei, et al.
Published: (2024)
by: Zhu, Lei, et al.
Published: (2024)
Generalizable Facial Expression Recognition
by: Zhang, Yuhang, et al.
Published: (2024)
by: Zhang, Yuhang, et al.
Published: (2024)
Segment Anything for Videos: A Systematic Survey
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
Efficient Egocentric Action Recognition with Multimodal Data
by: Calzavara, Marco, et al.
Published: (2025)
by: Calzavara, Marco, et al.
Published: (2025)
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
Similar Items
-
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
by: Wen, Yuhang, et al.
Published: (2023) -
HDBN: A Novel Hybrid Dual-branch Network for Robust Skeleton-based Action Recognition
by: Liu, Jinfu, et al.
Published: (2024) -
Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition
by: Liu, Mengyuan, et al.
Published: (2026) -
Exploring Explainability in Video Action Recognition
by: Saha, Avinab, et al.
Published: (2024) -
VG4D: Vision-Language Model Goes 4D Video Recognition
by: Deng, Zhichao, et al.
Published: (2024)