ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gupta, Rajan Das, Wei, Lei, Rahat, Md Yeasin, Fahad, Nafiz, Ahmed, Abir, Hui, Liew Tze
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915712091553792
author Gupta, Rajan Das
Wei, Lei
Rahat, Md Yeasin
Fahad, Nafiz
Ahmed, Abir
Hui, Liew Tze
author_facet Gupta, Rajan Das
Wei, Lei
Rahat, Md Yeasin
Fahad, Nafiz
Ahmed, Abir
Hui, Liew Tze
contents This study investigates the use of large language models (LLMs) for human behavior understanding by jointly leveraging motion and video data. We argue that integrating these complementary modalities is essential for capturing both fine-grained motion dynamics and contextual semantics of human actions, addressing the limitations of prior motion-only or video-only approaches. To this end, we propose ViMoNet, a multimodal vision-language framework trained through a two-stage alignment and instruction-tuning strategy that combines precise motion-text supervision with large-scale video-text data. We further introduce VIMOS, a multimodal dataset comprising human motion sequences, videos, and instruction-level annotations, along with ViMoNet-Bench, a standardized benchmark for evaluating behavior-centric reasoning. Experimental results demonstrate that ViMoNet consistently outperforms existing methods across caption generation, motion understanding, and human behavior interpretation tasks. The proposed framework shows significant potential in assistive healthcare applications, such as elderly monitoring, fall detection, and early identification of health risks in aging populations. This work contributes to the United Nations Sustainable Development Goal 3 (SDG 3: Good Health and Well-being) by enabling accessible AI-driven tools that promote universal health coverage, reduce preventable health issues, and enhance overall well-being.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09818
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video
Gupta, Rajan Das
Wei, Lei
Rahat, Md Yeasin
Fahad, Nafiz
Ahmed, Abir
Hui, Liew Tze
Computer Vision and Pattern Recognition
This study investigates the use of large language models (LLMs) for human behavior understanding by jointly leveraging motion and video data. We argue that integrating these complementary modalities is essential for capturing both fine-grained motion dynamics and contextual semantics of human actions, addressing the limitations of prior motion-only or video-only approaches. To this end, we propose ViMoNet, a multimodal vision-language framework trained through a two-stage alignment and instruction-tuning strategy that combines precise motion-text supervision with large-scale video-text data. We further introduce VIMOS, a multimodal dataset comprising human motion sequences, videos, and instruction-level annotations, along with ViMoNet-Bench, a standardized benchmark for evaluating behavior-centric reasoning. Experimental results demonstrate that ViMoNet consistently outperforms existing methods across caption generation, motion understanding, and human behavior interpretation tasks. The proposed framework shows significant potential in assistive healthcare applications, such as elderly monitoring, fall detection, and early identification of health risks in aging populations. This work contributes to the United Nations Sustainable Development Goal 3 (SDG 3: Good Health and Well-being) by enabling accessible AI-driven tools that promote universal health coverage, reduce preventable health issues, and enhance overall well-being.
title ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.09818