Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Shandilya, Utkarsh, Kappan, Marsha Mariya, Jain, Sanyam, Sharma, Vijeta |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention-Enhanced Lightweight Hourglass Network for Human Pose Estimation
by: Kappan, Marsha Mariya, et al.
Published: (2024)
by: Kappan, Marsha Mariya, et al.
Published: (2024)
LAPX: Lightweight Hourglass Network with Global Context
by: Zhao, Haopeng, et al.
Published: (2025)
by: Zhao, Haopeng, et al.
Published: (2025)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
by: Sharma, Saurav, et al.
Published: (2025)
by: Sharma, Saurav, et al.
Published: (2025)
Conformal Predictions for Human Action Recognition with Vision-Language Models
by: Tim, Bary, et al.
Published: (2025)
by: Tim, Bary, et al.
Published: (2025)
In Silico Pharmacokinetic and Molecular Docking Studies of Natural Plants against Essential Protein KRAS for Treatment of Pancreatic Cancer
by: Kappan, Marsha Mariya, et al.
Published: (2024)
by: Kappan, Marsha Mariya, et al.
Published: (2024)
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
by: Unmesh, Asim, et al.
Published: (2026)
by: Unmesh, Asim, et al.
Published: (2026)
DeepSeaNet: Improving Underwater Object Detection using EfficientDet
by: Jain, Sanyam
Published: (2023)
by: Jain, Sanyam
Published: (2023)
TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection
by: Chen, Hanning, et al.
Published: (2024)
by: Chen, Hanning, et al.
Published: (2024)
TimeSenCLIP: A Time Series Vision-Language Model for Remote Sensing
by: Jain, Pallavi, et al.
Published: (2025)
by: Jain, Pallavi, et al.
Published: (2025)
Adversarial Attack On Yolov5 For Traffic And Road Sign Detection
by: Jain, Sanyam
Published: (2023)
by: Jain, Sanyam
Published: (2023)
Pedestrian Attribute Recognition via CLIP based Prompt Vision-Language Fusion
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
Layout-Independent License Plate Recognition via Integrated Vision and Language Models
by: Shabaninia, Elham, et al.
Published: (2025)
by: Shabaninia, Elham, et al.
Published: (2025)
SMART-Vision: Survey of Modern Action Recognition Techniques in Vision
by: AlShami, Ali K., et al.
Published: (2025)
by: AlShami, Ali K., et al.
Published: (2025)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
by: Liu, Yating, et al.
Published: (2025)
by: Liu, Yating, et al.
Published: (2025)
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
by: Guruprasad, Pranav, et al.
Published: (2024)
by: Guruprasad, Pranav, et al.
Published: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
Vision-Language Models for Vision Tasks: A Survey
by: Zhang, Jingyi, et al.
Published: (2023)
by: Zhang, Jingyi, et al.
Published: (2023)
PE-CLIP: A Parameter-Efficient Fine-Tuning of Vision Language Models for Dynamic Facial Expression Recognition
by: Saadi, Ibtissam, et al.
Published: (2025)
by: Saadi, Ibtissam, et al.
Published: (2025)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Do Vision Models Develop Human-Like Progressive Difficulty Understanding?
by: Huang, Zeyi, et al.
Published: (2025)
by: Huang, Zeyi, et al.
Published: (2025)
An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
by: Silva, João Daniel, et al.
Published: (2025)
by: Silva, João Daniel, et al.
Published: (2025)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification
by: Addepalli, Sravanti, et al.
Published: (2023)
by: Addepalli, Sravanti, et al.
Published: (2023)
Generalized Recognition of Basic Surgical Actions Enables Skill Assessment and Vision-Language-Model-based Surgical Planning
by: Xu, Mengya, et al.
Published: (2026)
by: Xu, Mengya, et al.
Published: (2026)
Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments
by: Nagaonkar, Sankalp, et al.
Published: (2025)
by: Nagaonkar, Sankalp, et al.
Published: (2025)
Vision-Language Models Unlock Task-Centric Latent Actions
by: Nikulin, Alexander, et al.
Published: (2026)
by: Nikulin, Alexander, et al.
Published: (2026)
Resilience of Vision Transformers for Domain Generalisation in the Presence of Out-of-Distribution Noisy Images
by: Riaz, Hamza, et al.
Published: (2025)
by: Riaz, Hamza, et al.
Published: (2025)
CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
by: Imam, Raza, et al.
Published: (2024)
by: Imam, Raza, et al.
Published: (2024)
CLIP-VAD: Exploiting Vision-Language Models for Voice Activity Detection
by: Appiani, Andrea, et al.
Published: (2024)
by: Appiani, Andrea, et al.
Published: (2024)
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
by: Liu, Fan, et al.
Published: (2023)
by: Liu, Fan, et al.
Published: (2023)
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
by: Xiong, Zhitong, et al.
Published: (2025)
by: Xiong, Zhitong, et al.
Published: (2025)
Do Vision-Language Models Understand Compound Nouns?
by: Kumar, Sonal, et al.
Published: (2024)
by: Kumar, Sonal, et al.
Published: (2024)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
by: Liang, Wenqi, et al.
Published: (2025)
by: Liang, Wenqi, et al.
Published: (2025)
Joint Vision-Language Social Bias Removal for CLIP
by: Zhang, Haoyu, et al.
Published: (2024)
by: Zhang, Haoyu, et al.
Published: (2024)
FairCLIP: Harnessing Fairness in Vision-Language Learning
by: Luo, Yan, et al.
Published: (2024)
by: Luo, Yan, et al.
Published: (2024)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
by: Hu, Xiaoxing, et al.
Published: (2025)
by: Hu, Xiaoxing, et al.
Published: (2025)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis
by: Khurana, Mannat, et al.
Published: (2026)
by: Khurana, Mannat, et al.
Published: (2026)
Similar Items
-
Attention-Enhanced Lightweight Hourglass Network for Human Pose Estimation
by: Kappan, Marsha Mariya, et al.
Published: (2024) -
LAPX: Lightweight Hourglass Network with Global Context
by: Zhao, Haopeng, et al.
Published: (2025) -
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
by: Sharma, Saurav, et al.
Published: (2025) -
Conformal Predictions for Human Action Recognition with Vision-Language Models
by: Tim, Bary, et al.
Published: (2025) -
In Silico Pharmacokinetic and Molecular Docking Studies of Natural Plants against Essential Protein KRAS for Treatment of Pancreatic Cancer
by: Kappan, Marsha Mariya, et al.
Published: (2024)