The Potential and Limitations of Vision-Language Models for Human Motion Understanding: A Case Study in Data-Driven Stroke Rehabilitation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Victor, Kamalakannan, Naveenraj, Parnandi, Avinash, Schambra, Heidi, Fernandez-Granda, Carlos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Differentiable Biomechanics for Markerless Motion Capture in Upper Limb Stroke Rehabilitation: A Comparison with Optical Motion Capture
by: Unger, Tim, et al.
Published: (2024)
by: Unger, Tim, et al.
Published: (2024)
REST-HANDS: Rehabilitation with Egocentric Vision Using Smartglasses for Treatment of Hands after Surviving Stroke
by: Mucha, Wiktor, et al.
Published: (2024)
by: Mucha, Wiktor, et al.
Published: (2024)
Beyond Motion Pattern: An Empirical Study of Physical Forces for Human Motion Understanding
by: Dao, Anh, et al.
Published: (2025)
by: Dao, Anh, et al.
Published: (2025)
ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video
by: Gupta, Rajan Das, et al.
Published: (2025)
by: Gupta, Rajan Das, et al.
Published: (2025)
LS-GAN: Human Motion Synthesis with Latent-space GANs
by: Amballa, Avinash, et al.
Published: (2024)
by: Amballa, Avinash, et al.
Published: (2024)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025)
by: Hong, Wenyi, et al.
Published: (2025)
Exploring Vision Transformers for 3D Human Motion-Language Models with Motion Patches
by: Yu, Qing, et al.
Published: (2024)
by: Yu, Qing, et al.
Published: (2024)
Understanding differences in applying DETR to natural and medical images
by: Xu, Yanqi, et al.
Published: (2024)
by: Xu, Yanqi, et al.
Published: (2024)
MotionLLM: Understanding Human Behaviors from Human Motions and Videos
by: Chen, Ling-Hao, et al.
Published: (2024)
by: Chen, Ling-Hao, et al.
Published: (2024)
MMTA: Multi Membership Temporal Attention for Fine-Grained Stroke Rehabilitation Assessment
by: Helvaci, Halil Ismail, et al.
Published: (2026)
by: Helvaci, Halil Ismail, et al.
Published: (2026)
Deep Learning for Skeleton Based Human Motion Rehabilitation Assessment: A Benchmark
by: Ismail-Fawaz, Ali, et al.
Published: (2025)
by: Ismail-Fawaz, Ali, et al.
Published: (2025)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Encoder-Free Human Motion Understanding via Structured Motion Descriptions
by: Zhang, Yao, et al.
Published: (2026)
by: Zhang, Yao, et al.
Published: (2026)
The Limits of Learning from Pictures and Text: Vision-Language Models and Embodied Scene Understanding
by: Rosenberg, Gillian, et al.
Published: (2026)
by: Rosenberg, Gillian, et al.
Published: (2026)
Pushing the Limits of Vision-Language Models in Remote Sensing without Human Annotations
by: Cha, Keumgang, et al.
Published: (2024)
by: Cha, Keumgang, et al.
Published: (2024)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data
by: Li, Haoxin, et al.
Published: (2025)
by: Li, Haoxin, et al.
Published: (2025)
Hard Cases Detection in Motion Prediction by Vision-Language Foundation Models
by: Yang, Yi, et al.
Published: (2024)
by: Yang, Yi, et al.
Published: (2024)
Learning Priors of Human Motion With Vision Transformers
by: Falqueto, Placido, et al.
Published: (2025)
by: Falqueto, Placido, et al.
Published: (2025)
MotionFix: Text-Driven 3D Human Motion Editing
by: Athanasiou, Nikos, et al.
Published: (2024)
by: Athanasiou, Nikos, et al.
Published: (2024)
Temporal Visual Semantics-Induced Human Motion Understanding with Large Language Models
by: Xing, Zheng, et al.
Published: (2025)
by: Xing, Zheng, et al.
Published: (2025)
Moving by Looking: Towards Vision-Driven Avatar Motion Generation
by: Diomataris, Markos, et al.
Published: (2025)
by: Diomataris, Markos, et al.
Published: (2025)
Improving Large Vision-Language Models' Understanding for Flow Field Data
by: Zhang, Xiaomei, et al.
Published: (2025)
by: Zhang, Xiaomei, et al.
Published: (2025)
HVIS: A Human-like Vision and Inference System for Human Motion Prediction
by: Lyu, Kedi, et al.
Published: (2025)
by: Lyu, Kedi, et al.
Published: (2025)
HRTR: A Single-stage Transformer for Fine-grained Sub-second Action Segmentation in Stroke Rehabilitation
by: Helvaci, Halil Ismail, et al.
Published: (2025)
by: Helvaci, Halil Ismail, et al.
Published: (2025)
Unimotion: Unifying 3D Human Motion Synthesis and Understanding
by: Li, Chuqiao, et al.
Published: (2024)
by: Li, Chuqiao, et al.
Published: (2024)
Understanding the Transfer Limits of Vision Foundation Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
Multi-Resolution Generative Modeling of Human Motion from Limited Data
by: Moreno-Villamarín, David Eduardo, et al.
Published: (2024)
by: Moreno-Villamarín, David Eduardo, et al.
Published: (2024)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
by: Wang, Xinshun, et al.
Published: (2026)
by: Wang, Xinshun, et al.
Published: (2026)
MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition
by: Kim, Ye-eun, et al.
Published: (2025)
by: Kim, Ye-eun, et al.
Published: (2025)
The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
by: Verma, Arnav, et al.
Published: (2025)
by: Verma, Arnav, et al.
Published: (2025)
A Survey on Improving Human Robot Collaboration through Vision-and-Language Navigation
by: Yakolli, Nivedan, et al.
Published: (2025)
by: Yakolli, Nivedan, et al.
Published: (2025)
EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
by: Hou, Ruibing, et al.
Published: (2026)
by: Hou, Ruibing, et al.
Published: (2026)
Automatic Temporal Segmentation for Post-Stroke Rehabilitation: A Keypoint Detection and Temporal Segmentation Approach for Small Datasets
by: Lee, Jisoo, et al.
Published: (2025)
by: Lee, Jisoo, et al.
Published: (2025)
OrdinalBench: A Benchmark Dataset for Diagnosing Generalization Limits in Ordinal Number Understanding of Vision-Language Models
by: Tozaki, Yusuke, et al.
Published: (2026)
by: Tozaki, Yusuke, et al.
Published: (2026)
Semi-Supervised Masked Autoencoders: Unlocking Vision Transformer Potential with Limited Data
by: Faysal, Atik, et al.
Published: (2026)
by: Faysal, Atik, et al.
Published: (2026)
Understanding Degradation with Vision Language Model
by: Lan, Guanzhou, et al.
Published: (2026)
by: Lan, Guanzhou, et al.
Published: (2026)
Similar Items
-
Differentiable Biomechanics for Markerless Motion Capture in Upper Limb Stroke Rehabilitation: A Comparison with Optical Motion Capture
by: Unger, Tim, et al.
Published: (2024) -
REST-HANDS: Rehabilitation with Egocentric Vision Using Smartglasses for Treatment of Hands after Surviving Stroke
by: Mucha, Wiktor, et al.
Published: (2024) -
Beyond Motion Pattern: An Empirical Study of Physical Forces for Human Motion Understanding
by: Dao, Anh, et al.
Published: (2025) -
ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video
by: Gupta, Rajan Das, et al.
Published: (2025) -
LS-GAN: Human Motion Synthesis with Latent-space GANs
by: Amballa, Avinash, et al.
Published: (2024)