Saved in:
Bibliographic Details
Main Authors: Wu, Bizhu, Xie, Jinheng, Shen, Keming, Kong, Zhe, Ren, Jianfeng, Bai, Ruibin, Qu, Rong, Shen, Linlin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.02478
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917975704993792
author Wu, Bizhu
Xie, Jinheng
Shen, Keming
Kong, Zhe
Ren, Jianfeng
Bai, Ruibin
Qu, Rong
Shen, Linlin
author_facet Wu, Bizhu
Xie, Jinheng
Shen, Keming
Kong, Zhe
Ren, Jianfeng
Bai, Ruibin
Qu, Rong
Shen, Linlin
contents Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM
format Preprint
id arxiv_https___arxiv_org_abs_2504_02478
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities
Wu, Bizhu
Xie, Jinheng
Shen, Keming
Kong, Zhe
Ren, Jianfeng
Bai, Ruibin
Qu, Rong
Shen, Linlin
Computer Vision and Pattern Recognition
Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM
title MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.02478