Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, MingZe, Kazi, Madiha
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916793341181952
author Tang, MingZe
Kazi, Madiha
author_facet Tang, MingZe
Kazi, Madiha
contents This study explores human action recognition using a three-class subset of the COCO image corpus, benchmarking models from simple fully connected networks to transformer architectures. The binary Vision Transformer (ViT) achieved 90% mean test accuracy, significantly exceeding multiclass classifiers such as convolutional networks (approximately 35%) and CLIP-based models (approximately 62-64%). A one-way ANOVA (F = 61.37, p < 0.001) confirmed these differences are statistically significant. Qualitative analysis with SHAP explainer and LeGrad heatmaps indicated that the ViT localizes pose-specific regions (e.g., lower limbs for walking or running), while simpler feed-forward models often focus on background textures, explaining their errors. These findings emphasize the data efficiency of transformer representations and the importance of explainability techniques in diagnosing class-specific failures.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11678
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
Tang, MingZe
Kazi, Madiha
Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.0
This study explores human action recognition using a three-class subset of the COCO image corpus, benchmarking models from simple fully connected networks to transformer architectures. The binary Vision Transformer (ViT) achieved 90% mean test accuracy, significantly exceeding multiclass classifiers such as convolutional networks (approximately 35%) and CLIP-based models (approximately 62-64%). A one-way ANOVA (F = 61.37, p < 0.001) confirmed these differences are statistically significant. Qualitative analysis with SHAP explainer and LeGrad heatmaps indicated that the ViT localizes pose-specific regions (e.g., lower limbs for walking or running), while simpler feed-forward models often focus on background textures, explaining their errors. These findings emphasize the data efficiency of transformer representations and the importance of explainability techniques in diagnosing class-specific failures.
title Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
topic Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.0
url https://arxiv.org/abs/2506.11678