Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Abid, Abderrazek, Ho, Thanh-Cong, Karray, Fakhri
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911269531942912
author Abid, Abderrazek
Ho, Thanh-Cong
Karray, Fakhri
author_facet Abid, Abderrazek
Ho, Thanh-Cong
Karray, Fakhri
contents As generative AI continues to evolve, Vision Language Models (VLMs) have emerged as promising tools in various healthcare applications. One area that remains relatively underexplored is their use in human activity recognition (HAR) for remote health monitoring. VLMs offer notable strengths, including greater flexibility and the ability to overcome some of the constraints of traditional deep learning models. However, a key challenge in applying VLMs to HAR lies in the difficulty of evaluating their dynamic and often non-deterministic outputs. To address this gap, we introduce a descriptive caption data set and propose comprehensive evaluation methods to evaluate VLMs in HAR. Through comparative experiments with state-of-the-art deep learning models, our findings demonstrate that VLMs achieve comparable performance and, in some cases, even surpass conventional approaches in terms of accuracy. This work contributes a strong benchmark and opens new possibilities for the integration of VLMs into intelligent healthcare systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21424
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
Abid, Abderrazek
Ho, Thanh-Cong
Karray, Fakhri
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
As generative AI continues to evolve, Vision Language Models (VLMs) have emerged as promising tools in various healthcare applications. One area that remains relatively underexplored is their use in human activity recognition (HAR) for remote health monitoring. VLMs offer notable strengths, including greater flexibility and the ability to overcome some of the constraints of traditional deep learning models. However, a key challenge in applying VLMs to HAR lies in the difficulty of evaluating their dynamic and often non-deterministic outputs. To address this gap, we introduce a descriptive caption data set and propose comprehensive evaluation methods to evaluate VLMs in HAR. Through comparative experiments with state-of-the-art deep learning models, our findings demonstrate that VLMs achieve comparable performance and, in some cases, even surpass conventional approaches in terms of accuracy. This work contributes a strong benchmark and opens new possibilities for the integration of VLMs into intelligent healthcare systems.
title Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2510.21424