Human Motion Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913758724489216 |
|---|---|
| author | Li, Lei Jia, Sen Wang, Jianhao Jiang, Zhongyu Zhou, Feng Dai, Ju Zhang, Tianfang Wu, Zongkai Hwang, Jenq-Neng |
| author_facet | Li, Lei Jia, Sen Wang, Jianhao Jiang, Zhongyu Zhou, Feng Dai, Ju Zhang, Tianfang Wu, Zongkai Hwang, Jenq-Neng |
| contents | This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model's ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Our code and models are available on the project website: https://github.com/ILGLJ/LLaMo. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_16805 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Human Motion Instruction Tuning Li, Lei Jia, Sen Wang, Jianhao Jiang, Zhongyu Zhou, Feng Dai, Ju Zhang, Tianfang Wu, Zongkai Hwang, Jenq-Neng Artificial Intelligence Computer Vision and Pattern Recognition This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model's ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Our code and models are available on the project website: https://github.com/ILGLJ/LLaMo. |
| title | Human Motion Instruction Tuning |
| topic | Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2411.16805 |