OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909973420703744 |
|---|---|
| author | Bu, Wendong Pan, Kaihang Lin, Yuze Li, Jiacheng Shen, Kai Zhang, Wenqiao Li, Juncheng Xiao, Jun Tang, Siliang |
| author_facet | Bu, Wendong Pan, Kaihang Lin, Yuze Li, Jiacheng Shen, Kai Zhang, Wenqiao Li, Juncheng Xiao, Jun Tang, Siliang |
| contents | Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for free-form and omni-objective generation. To address this, we propose OmniMoGen, a unified framework that enables versatile motion generation through interleaved text-motion instructions. Built upon a concise RVQ-VAE and transformer architecture, OmniMoGen supports end-to-end instruction-driven motion generation. We construct X2Mo, a large-scale dataset of over 137K interleaved text-motion instructions, and introduce AnyContext, a benchmark for evaluating interleaved motion generation. Experiments show that OmniMoGen achieves state-of-the-art performance on text-to-motion, motion editing, and AnyContext, exhibiting emerging capabilities such as compositional editing, self-reflective generation, and knowledge-informed generation. These results mark a step toward the next intelligent motion generation. Project Page: https://OmniMoGen.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_19159 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions Bu, Wendong Pan, Kaihang Lin, Yuze Li, Jiacheng Shen, Kai Zhang, Wenqiao Li, Juncheng Xiao, Jun Tang, Siliang Computer Vision and Pattern Recognition Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for free-form and omni-objective generation. To address this, we propose OmniMoGen, a unified framework that enables versatile motion generation through interleaved text-motion instructions. Built upon a concise RVQ-VAE and transformer architecture, OmniMoGen supports end-to-end instruction-driven motion generation. We construct X2Mo, a large-scale dataset of over 137K interleaved text-motion instructions, and introduce AnyContext, a benchmark for evaluating interleaved motion generation. Experiments show that OmniMoGen achieves state-of-the-art performance on text-to-motion, motion editing, and AnyContext, exhibiting emerging capabilities such as compositional editing, self-reflective generation, and knowledge-informed generation. These results mark a step toward the next intelligent motion generation. Project Page: https://OmniMoGen.github.io/. |
| title | OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.19159 |