Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, MingZe, Kazi, Madiha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Context-Aware Full Body Anonymization using Text-to-Image Diffusion Models
by: Zwick, Pascal, et al.
Published: (2024)
by: Zwick, Pascal, et al.
Published: (2024)
Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
by: Tangtartharakul, Gene, et al.
Published: (2026)
by: Tangtartharakul, Gene, et al.
Published: (2026)
EncQA: Benchmarking Vision-Language Models on Visual Encodings for Charts
by: Mukherjee, Kushin, et al.
Published: (2025)
by: Mukherjee, Kushin, et al.
Published: (2025)
Measuring proximity to standard planes during fetal brain ultrasound scanning
by: Di Vece, Chiara, et al.
Published: (2024)
by: Di Vece, Chiara, et al.
Published: (2024)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)
by: He, Mengqi, et al.
Published: (2025)
Exploring Transfer Learning for Deep Learning Polyp Detection in Colonoscopy Images Using YOLOv8
by: Vazquez, Fabian, et al.
Published: (2025)
by: Vazquez, Fabian, et al.
Published: (2025)
CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
by: Li, Jiaxuan, et al.
Published: (2025)
by: Li, Jiaxuan, et al.
Published: (2025)
SeNeDiF-OOD: Semantic Nested Dichotomy Fusion for Out-of-Distribution Detection Methodology in Open-World Classification. A Case Study on Monument Style Classification
by: Antequera-Sánchez, Ignacio, et al.
Published: (2026)
by: Antequera-Sánchez, Ignacio, et al.
Published: (2026)
Deep EM with Hierarchical Latent Label Modelling for Multi-Site Prostate Lesion Segmentation
by: Yan, Wen, et al.
Published: (2026)
by: Yan, Wen, et al.
Published: (2026)
Quaternion Convolutional Neural Networks: Current Advances and Future Directions
by: Altamirano-Gomez, Gerardo, et al.
Published: (2023)
by: Altamirano-Gomez, Gerardo, et al.
Published: (2023)
Efficient Diffusion Training through Parallelization with Truncated Karhunen-Loève Expansion
by: Ren, Yumeng, et al.
Published: (2025)
by: Ren, Yumeng, et al.
Published: (2025)
A Review of Pseudo-Labeling for Computer Vision
by: Kage, Patrick, et al.
Published: (2024)
by: Kage, Patrick, et al.
Published: (2024)
Skeleton-based sign language recognition using a dual-stream spatio-temporal dynamic graph convolutional network
by: Liu, Liangjin, et al.
Published: (2025)
by: Liu, Liangjin, et al.
Published: (2025)
MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
by: Liu, Ying, et al.
Published: (2025)
by: Liu, Ying, et al.
Published: (2025)
MAR-MAER: Metric-Aware and Ambiguity-Adaptive Autoregressive Image Generation
by: Dong, Kai, et al.
Published: (2026)
by: Dong, Kai, et al.
Published: (2026)
Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation
by: Kuprashevich, Maksim, et al.
Published: (2024)
by: Kuprashevich, Maksim, et al.
Published: (2024)
Synthetic Image Generation in Cyber Influence Operations: An Emergent Threat?
by: Mathys, Melanie, et al.
Published: (2024)
by: Mathys, Melanie, et al.
Published: (2024)
FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare
by: Lekadir, Karim, et al.
Published: (2023)
by: Lekadir, Karim, et al.
Published: (2023)
Language as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity
by: Tang, MingZe, et al.
Published: (2025)
by: Tang, MingZe, et al.
Published: (2025)
CerberusDet: Unified Multi-Dataset Object Detection
by: Tolstykh, Irina, et al.
Published: (2024)
by: Tolstykh, Irina, et al.
Published: (2024)
Comparing Zealous and Restrained AI Recommendations in a Real-World Human-AI Collaboration Task
by: Xu, Chengyuan, et al.
Published: (2024)
by: Xu, Chengyuan, et al.
Published: (2024)
Generation of Complex 3D Human Motion by Temporal and Spatial Composition of Diffusion Models
by: Mandelli, Lorenzo, et al.
Published: (2024)
by: Mandelli, Lorenzo, et al.
Published: (2024)
Unified Local and Global Attention Interaction Modeling for Vision Transformers
by: Nguyen, Tan, et al.
Published: (2024)
by: Nguyen, Tan, et al.
Published: (2024)
VACoDe: Visual Augmented Contrastive Decoding
by: Kim, Sihyeon, et al.
Published: (2024)
by: Kim, Sihyeon, et al.
Published: (2024)
Towards Infusing Auxiliary Knowledge for Distracted Driver Detection
by: Balappanawar, Ishwar B, et al.
Published: (2024)
by: Balappanawar, Ishwar B, et al.
Published: (2024)
Parameterizing Dataset Distillation via Gaussian Splatting
by: Jiang, Chenyang, et al.
Published: (2025)
by: Jiang, Chenyang, et al.
Published: (2025)
Vision Transformer-based Model for Severity Quantification of Lung Pneumonia Using Chest X-ray Images
by: Slika, Bouthaina, et al.
Published: (2023)
by: Slika, Bouthaina, et al.
Published: (2023)
Universal Adversarial Perturbations for Vision-Language Pre-trained Models
by: Zhang, Peng-Fei, et al.
Published: (2024)
by: Zhang, Peng-Fei, et al.
Published: (2024)
From CNNs to Transformers in Multimodal Human Action Recognition: A Survey
by: Shaikh, Muhammad Bilal, et al.
Published: (2024)
by: Shaikh, Muhammad Bilal, et al.
Published: (2024)
Domain Generalized Stereo Matching with Uncertainty-guided Data Augmentation
by: Du, Shuangli, et al.
Published: (2025)
by: Du, Shuangli, et al.
Published: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
by: Rudman, William, et al.
Published: (2026)
by: Rudman, William, et al.
Published: (2026)
Visual Language Models show widespread visual deficits on neuropsychological tests
by: Tangtartharakul, Gene, et al.
Published: (2025)
by: Tangtartharakul, Gene, et al.
Published: (2025)
Lookism: The overlooked bias in computer vision
by: Gulati, Aditya, et al.
Published: (2024)
by: Gulati, Aditya, et al.
Published: (2024)
Scene-wise Adaptive Network for Dynamic Cold-start Scenes Optimization in CTR Prediction
by: Li, Wenhao, et al.
Published: (2024)
by: Li, Wenhao, et al.
Published: (2024)
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
by: Bouzid, Hamza, et al.
Published: (2023)
by: Bouzid, Hamza, et al.
Published: (2023)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
by: Nemitz, Jonathan, et al.
Published: (2026)
by: Nemitz, Jonathan, et al.
Published: (2026)
Synthetic Photography Detection: A Visual Guidance for Identifying Synthetic Images Created by AI
by: Mathys, Melanie, et al.
Published: (2024)
by: Mathys, Melanie, et al.
Published: (2024)
Surrealistic-like Image Generation with Vision-Language Models
by: Ayten, Elif, et al.
Published: (2024)
by: Ayten, Elif, et al.
Published: (2024)
Sparse vs Contiguous Adversarial Pixel Perturbations in Multimodal Models: An Empirical Analysis
by: Botocan, Cristian-Alexandru, et al.
Published: (2024)
by: Botocan, Cristian-Alexandru, et al.
Published: (2024)
Similar Items
-
Context-Aware Full Body Anonymization using Text-to-Image Diffusion Models
by: Zwick, Pascal, et al.
Published: (2024) -
Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
by: Tangtartharakul, Gene, et al.
Published: (2026) -
EncQA: Benchmarking Vision-Language Models on Visual Encodings for Charts
by: Mukherjee, Kushin, et al.
Published: (2025) -
Measuring proximity to standard planes during fetal brain ultrasound scanning
by: Di Vece, Chiara, et al.
Published: (2024) -
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)