Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.02478 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917975704993792 |
|---|---|
| author | Wu, Bizhu Xie, Jinheng Shen, Keming Kong, Zhe Ren, Jianfeng Bai, Ruibin Qu, Rong Shen, Linlin |
| author_facet | Wu, Bizhu Xie, Jinheng Shen, Keming Kong, Zhe Ren, Jianfeng Bai, Ruibin Qu, Rong Shen, Linlin |
| contents | Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_02478 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities Wu, Bizhu Xie, Jinheng Shen, Keming Kong, Zhe Ren, Jianfeng Bai, Ruibin Qu, Rong Shen, Linlin Computer Vision and Pattern Recognition Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM |
| title | MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.02478 |