Are Visual-Language Models Effective in Action Recognition? A Comparative Study
Fuente:
arXiv
Saved in:
| Main Authors: | Ali, Mahmoud, Yang, Di, Brémond, François |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AM Flow: Adapters for Temporal Processing in Action Recognition
by: Agrawal, Tanay, et al.
Published: (2024)
by: Agrawal, Tanay, et al.
Published: (2024)
LAC: Latent Action Composition for Skeleton-based Action Segmentation
by: Yang, Di, et al.
Published: (2023)
by: Yang, Di, et al.
Published: (2023)
Temporally Propagated Masks and Bounding Boxes: Combining the Best of Both Worlds for Multi-Object Tracking
by: Stanczyk, Tomasz, et al.
Published: (2024)
by: Stanczyk, Tomasz, et al.
Published: (2024)
B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition
by: Poddar, Nishit, et al.
Published: (2026)
by: Poddar, Nishit, et al.
Published: (2026)
Introducing Gating and Context into Temporal Action Detection
by: Reka, Aglind, et al.
Published: (2024)
by: Reka, Aglind, et al.
Published: (2024)
Weakly-supervised Autism Severity Assessment in Long Videos
by: Ali, Abid, et al.
Published: (2024)
by: Ali, Abid, et al.
Published: (2024)
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
No Train Yet Gain: Towards Generic Multi-Object Tracking in Sports and Beyond
by: Stanczyk, Tomasz, et al.
Published: (2025)
by: Stanczyk, Tomasz, et al.
Published: (2025)
Loose Social-Interaction Recognition in Real-world Therapy Scenarios
by: Ali, Abid, et al.
Published: (2024)
by: Ali, Abid, et al.
Published: (2024)
LIA-X: Interpretable Latent Portrait Animator
by: Wang, Yaohui, et al.
Published: (2025)
by: Wang, Yaohui, et al.
Published: (2025)
What Matters in Autonomous Driving Anomaly Detection: A Weakly Supervised Horizon
by: Tiwari, Utkarsh, et al.
Published: (2024)
by: Tiwari, Utkarsh, et al.
Published: (2024)
MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition
by: Kassab, Hozaifa, et al.
Published: (2024)
by: Kassab, Hozaifa, et al.
Published: (2024)
MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
by: Zeeshan, Muhammad Osama, et al.
Published: (2025)
by: Zeeshan, Muhammad Osama, et al.
Published: (2025)
An Effective End-to-End Solution for Multimodal Action Recognition
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
SBF: An Effective Representation to Augment Skeleton for Video-based Human Action Recognition
by: Peng, Zhuoxuan, et al.
Published: (2026)
by: Peng, Zhuoxuan, et al.
Published: (2026)
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
by: Sun, Baoli, et al.
Published: (2025)
by: Sun, Baoli, et al.
Published: (2025)
EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis
by: Cha, Junuk, et al.
Published: (2025)
by: Cha, Junuk, et al.
Published: (2025)
Zero-Shot Skeleton-based Action Recognition with Dual Visual-Text Alignment
by: Kuang, Jidong, et al.
Published: (2024)
by: Kuang, Jidong, et al.
Published: (2024)
A Comparative Study of Continuous Sign Language Recognition Techniques
by: Alyami, Sarah, et al.
Published: (2024)
by: Alyami, Sarah, et al.
Published: (2024)
TIM: A Time Interval Machine for Audio-Visual Action Recognition
by: Chalk, Jacob, et al.
Published: (2024)
by: Chalk, Jacob, et al.
Published: (2024)
Multimodal Large Models Are Effective Action Anticipators
by: Wang, Binglu, et al.
Published: (2025)
by: Wang, Binglu, et al.
Published: (2025)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
ScenarioCLIP: Pretrained Transferable Visual Language Models and Action-Genome Dataset for Natural Scene Analysis
by: Sinha, Advik, et al.
Published: (2025)
by: Sinha, Advik, et al.
Published: (2025)
Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs
by: Orjuela, Daniel Yezid Guarnizo, et al.
Published: (2026)
by: Orjuela, Daniel Yezid Guarnizo, et al.
Published: (2026)
Advancing Human Action Recognition with Foundation Models trained on Unlabeled Public Videos
by: Qian, Yang, et al.
Published: (2024)
by: Qian, Yang, et al.
Published: (2024)
A Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
by: Gao, Mingjian, et al.
Published: (2026)
by: Gao, Mingjian, et al.
Published: (2026)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Towards Unified Facial Action Unit Recognition Framework by Large Language Models
by: Hu, Guohong, et al.
Published: (2024)
by: Hu, Guohong, et al.
Published: (2024)
SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
Noise-Tolerant Learning for Audio-Visual Action Recognition
by: Han, Haochen, et al.
Published: (2022)
by: Han, Haochen, et al.
Published: (2022)
Democratizing Fine-grained Visual Recognition with Large Language Models
by: Liu, Mingxuan, et al.
Published: (2024)
by: Liu, Mingxuan, et al.
Published: (2024)
Prompting Visual-Language Models for Dynamic Facial Expression Recognition
by: Zhao, Zengqun, et al.
Published: (2023)
by: Zhao, Zengqun, et al.
Published: (2023)
A Comprehensive Survey of Masked Faces: Recognition, Detection, and Unmasking
by: Mahmoud, Mohamed, et al.
Published: (2024)
by: Mahmoud, Mohamed, et al.
Published: (2024)
Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition
by: Qian, Zefeng, et al.
Published: (2025)
by: Qian, Zefeng, et al.
Published: (2025)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
A Comprehensive Review of Few-shot Action Recognition
by: Wanyan, Yuyang, et al.
Published: (2024)
by: Wanyan, Yuyang, et al.
Published: (2024)
LLM Enhanced Action Recognition via Hierarchical Global-Local Skeleton-Language Model
by: Wang, Ruosi, et al.
Published: (2026)
by: Wang, Ruosi, et al.
Published: (2026)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
Similar Items
-
AM Flow: Adapters for Temporal Processing in Action Recognition
by: Agrawal, Tanay, et al.
Published: (2024) -
LAC: Latent Action Composition for Skeleton-based Action Segmentation
by: Yang, Di, et al.
Published: (2023) -
Temporally Propagated Masks and Bounding Boxes: Combining the Best of Both Worlds for Multi-Object Tracking
by: Stanczyk, Tomasz, et al.
Published: (2024) -
B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition
by: Poddar, Nishit, et al.
Published: (2026) -
Introducing Gating and Context into Temporal Action Detection
by: Reka, Aglind, et al.
Published: (2024)