RNNs, CNNs and Transformers in Human Action Recognition: A Survey and a Hybrid Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Alomar, Khaled, Aysel, Halil Ibrahim, Cai, Xiaohao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Semantic Segmentation by Semantic Proportions
por: Aysel, Halil Ibrahim, et al.
Publicado: (2023)
por: Aysel, Halil Ibrahim, et al.
Publicado: (2023)
VORTEX: Challenging CNNs at Texture Recognition by using Vision Transformers with Orderless and Randomized Token Encodings
por: Scabini, Leonardo, et al.
Publicado: (2025)
por: Scabini, Leonardo, et al.
Publicado: (2025)
TraNCE: Transformative Non-linear Concept Explainer for CNNs
por: Akpudo, Ugochukwu Ejike, et al.
Publicado: (2025)
por: Akpudo, Ugochukwu Ejike, et al.
Publicado: (2025)
Real-Time Human Action Recognition on Embedded Platforms
por: Wang, Ruiqi, et al.
Publicado: (2024)
por: Wang, Ruiqi, et al.
Publicado: (2024)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2024)
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2024)
Transformer-based Models to Deal with Heterogeneous Environments in Human Activity Recognition
por: EK, Sannara, et al.
Publicado: (2022)
por: EK, Sannara, et al.
Publicado: (2022)
Do Transformers Understand Ancient Roman Coin Motifs Better than CNNs?
por: Reid, David, et al.
Publicado: (2026)
por: Reid, David, et al.
Publicado: (2026)
Concept-Based Explainable Artificial Intelligence: Metrics and Benchmarks
por: Aysel, Halil Ibrahim, et al.
Publicado: (2025)
por: Aysel, Halil Ibrahim, et al.
Publicado: (2025)
Explaining Model Overfitting in CNNs via GMM Clustering
por: Dou, Hui, et al.
Publicado: (2024)
por: Dou, Hui, et al.
Publicado: (2024)
Survey of Action Recognition, Spotting and Spatio-Temporal Localization in Soccer -- Current Trends and Research Perspectives
por: Seweryn, Karolina, et al.
Publicado: (2023)
por: Seweryn, Karolina, et al.
Publicado: (2023)
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
por: Babey, Nicholas, et al.
Publicado: (2025)
por: Babey, Nicholas, et al.
Publicado: (2025)
Efficient Hyperparameter Importance Assessment for CNNs
por: Wang, Ruinan, et al.
Publicado: (2024)
por: Wang, Ruinan, et al.
Publicado: (2024)
Detecção da Psoríase Utilizando Visão Computacional: Uma Abordagem Comparativa Entre CNNs e Vision Transformers
por: Lucena, Natanael, et al.
Publicado: (2025)
por: Lucena, Natanael, et al.
Publicado: (2025)
CNNs Avoid Curse of Dimensionality by Learning on Patches
por: Madala, Vamshi C., et al.
Publicado: (2022)
por: Madala, Vamshi C., et al.
Publicado: (2022)
Explaning with trees: interpreting CNNs using hierarchies
por: Rodrigues, Caroline Mazini, et al.
Publicado: (2024)
por: Rodrigues, Caroline Mazini, et al.
Publicado: (2024)
From Ground to Air: Noise Robustness in Vision Transformers and CNNs for Event-Based Vehicle Classification with Potential UAV Applications
por: Almesafri, Nouf, et al.
Publicado: (2025)
por: Almesafri, Nouf, et al.
Publicado: (2025)
Feature Hallucination for Self-supervised Action Recognition
por: Wang, Lei, et al.
Publicado: (2025)
por: Wang, Lei, et al.
Publicado: (2025)
Evolving Skeletons: Motion Dynamics in Action Recognition
por: Qiu, Jushang, et al.
Publicado: (2025)
por: Qiu, Jushang, et al.
Publicado: (2025)
SkelVIT: Consensus of Vision Transformers for a Lightweight Skeleton-Based Action Recognition System
por: Karadag, Ozge Oztimur
Publicado: (2023)
por: Karadag, Ozge Oztimur
Publicado: (2023)
Evaluating the Stability of Semantic Concept Representations in CNNs for Robust Explainability
por: Mikriukov, Georgii, et al.
Publicado: (2023)
por: Mikriukov, Georgii, et al.
Publicado: (2023)
Hybrid Training for Vision-Language-Action Models
por: Mazzaglia, Pietro, et al.
Publicado: (2025)
por: Mazzaglia, Pietro, et al.
Publicado: (2025)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
por: Tang, Haotian, et al.
Publicado: (2024)
por: Tang, Haotian, et al.
Publicado: (2024)
Reliable Evaluation of Attribution Maps in CNNs: A Perturbation-Based Approach
por: Nieradzik, Lars, et al.
Publicado: (2024)
por: Nieradzik, Lars, et al.
Publicado: (2024)
Stacked Ensemble of Fine-Tuned CNNs for Knee Osteoarthritis Severity Grading
por: Gupta, Adarsh, et al.
Publicado: (2025)
por: Gupta, Adarsh, et al.
Publicado: (2025)
SkelMamba: A State Space Model for Efficient Skeleton Action Recognition of Neurological Disorders
por: Martinel, Niki, et al.
Publicado: (2024)
por: Martinel, Niki, et al.
Publicado: (2024)
Skeleton-based Action Recognition with Non-linear Dependency Modeling and Hilbert-Schmidt Independence Criterion
por: Yang, Yuheng
Publicado: (2024)
por: Yang, Yuheng
Publicado: (2024)
A Survey on Efficient Vision-Language-Action Models
por: Yu, Zhaoshu, et al.
Publicado: (2025)
por: Yu, Zhaoshu, et al.
Publicado: (2025)
A Comparative Study of Custom CNNs, Pre-trained Models, and Transfer Learning Across Multiple Visual Datasets
por: Akhand, Annoor Sharara
Publicado: (2026)
por: Akhand, Annoor Sharara
Publicado: (2026)
Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
por: Tang, MingZe, et al.
Publicado: (2025)
por: Tang, MingZe, et al.
Publicado: (2025)
Hypergraph-based Multi-View Action Recognition using Event Cameras
por: Gao, Yue, et al.
Publicado: (2024)
por: Gao, Yue, et al.
Publicado: (2024)
One-Frame Calibration with Siamese Network in Facial Action Unit Recognition
por: Feng, Shuangquan, et al.
Publicado: (2024)
por: Feng, Shuangquan, et al.
Publicado: (2024)
EITNet: An IoT-Enhanced Framework for Real-Time Basketball Action Recognition
por: Liu, Jingyu, et al.
Publicado: (2024)
por: Liu, Jingyu, et al.
Publicado: (2024)
Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation
por: Li, Wenhao, et al.
Publicado: (2023)
por: Li, Wenhao, et al.
Publicado: (2023)
Federated Learning for Video Violence Detection: Complementary Roles of Lightweight CNNs and Vision-Language Models for Energy-Efficient Use
por: Thuau, Sébastien, et al.
Publicado: (2025)
por: Thuau, Sébastien, et al.
Publicado: (2025)
Multimodal Attack Detection for Action Recognition Models
por: Mumcu, Furkan, et al.
Publicado: (2024)
por: Mumcu, Furkan, et al.
Publicado: (2024)
Advanced Arabic Alphabet Sign Language Recognition Using Transfer Learning and Transformer Models
por: Balat, Mazen, et al.
Publicado: (2024)
por: Balat, Mazen, et al.
Publicado: (2024)
A Survey of Deep Learning for Group-level Emotion Recognition
por: Huang, Xiaohua, et al.
Publicado: (2024)
por: Huang, Xiaohua, et al.
Publicado: (2024)
Beyond Conventional Transformers: The Medical X-ray Attention (MXA) Block for Improved Multi-Label Diagnosis Using Knowledge Distillation
por: Rand, Amit, et al.
Publicado: (2025)
por: Rand, Amit, et al.
Publicado: (2025)
Bird Eye-View to Street-View: A Survey
por: Bajbaa, Khawlah, et al.
Publicado: (2024)
por: Bajbaa, Khawlah, et al.
Publicado: (2024)
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
por: Lin, Nie, et al.
Publicado: (2025)
por: Lin, Nie, et al.
Publicado: (2025)
Ejemplares similares
-
Semantic Segmentation by Semantic Proportions
por: Aysel, Halil Ibrahim, et al.
Publicado: (2023) -
VORTEX: Challenging CNNs at Texture Recognition by using Vision Transformers with Orderless and Randomized Token Encodings
por: Scabini, Leonardo, et al.
Publicado: (2025) -
TraNCE: Transformative Non-linear Concept Explainer for CNNs
por: Akpudo, Ugochukwu Ejike, et al.
Publicado: (2025) -
Real-Time Human Action Recognition on Embedded Platforms
por: Wang, Ruiqi, et al.
Publicado: (2024) -
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2024)