Human Motion Instruction Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Lei, Jia, Sen, Wang, Jianhao, Jiang, Zhongyu, Zhou, Feng, Dai, Ju, Zhang, Tianfang, Wu, Zongkai, Hwang, Jenq-Neng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913758724489216
author Li, Lei
Jia, Sen
Wang, Jianhao
Jiang, Zhongyu
Zhou, Feng
Dai, Ju
Zhang, Tianfang
Wu, Zongkai
Hwang, Jenq-Neng
author_facet Li, Lei
Jia, Sen
Wang, Jianhao
Jiang, Zhongyu
Zhou, Feng
Dai, Ju
Zhang, Tianfang
Wu, Zongkai
Hwang, Jenq-Neng
contents This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model's ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Our code and models are available on the project website: https://github.com/ILGLJ/LLaMo.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16805
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Human Motion Instruction Tuning
Li, Lei
Jia, Sen
Wang, Jianhao
Jiang, Zhongyu
Zhou, Feng
Dai, Ju
Zhang, Tianfang
Wu, Zongkai
Hwang, Jenq-Neng
Artificial Intelligence
Computer Vision and Pattern Recognition
This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model's ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Our code and models are available on the project website: https://github.com/ILGLJ/LLaMo.
title Human Motion Instruction Tuning
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.16805