OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bu, Wendong, Pan, Kaihang, Lin, Yuze, Li, Jiacheng, Shen, Kai, Zhang, Wenqiao, Li, Juncheng, Xiao, Jun, Tang, Siliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909973420703744
author Bu, Wendong
Pan, Kaihang
Lin, Yuze
Li, Jiacheng
Shen, Kai
Zhang, Wenqiao
Li, Juncheng
Xiao, Jun
Tang, Siliang
author_facet Bu, Wendong
Pan, Kaihang
Lin, Yuze
Li, Jiacheng
Shen, Kai
Zhang, Wenqiao
Li, Juncheng
Xiao, Jun
Tang, Siliang
contents Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for free-form and omni-objective generation. To address this, we propose OmniMoGen, a unified framework that enables versatile motion generation through interleaved text-motion instructions. Built upon a concise RVQ-VAE and transformer architecture, OmniMoGen supports end-to-end instruction-driven motion generation. We construct X2Mo, a large-scale dataset of over 137K interleaved text-motion instructions, and introduce AnyContext, a benchmark for evaluating interleaved motion generation. Experiments show that OmniMoGen achieves state-of-the-art performance on text-to-motion, motion editing, and AnyContext, exhibiting emerging capabilities such as compositional editing, self-reflective generation, and knowledge-informed generation. These results mark a step toward the next intelligent motion generation. Project Page: https://OmniMoGen.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19159
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
Bu, Wendong
Pan, Kaihang
Lin, Yuze
Li, Jiacheng
Shen, Kai
Zhang, Wenqiao
Li, Juncheng
Xiao, Jun
Tang, Siliang
Computer Vision and Pattern Recognition
Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for free-form and omni-objective generation. To address this, we propose OmniMoGen, a unified framework that enables versatile motion generation through interleaved text-motion instructions. Built upon a concise RVQ-VAE and transformer architecture, OmniMoGen supports end-to-end instruction-driven motion generation. We construct X2Mo, a large-scale dataset of over 137K interleaved text-motion instructions, and introduce AnyContext, a benchmark for evaluating interleaved motion generation. Experiments show that OmniMoGen achieves state-of-the-art performance on text-to-motion, motion editing, and AnyContext, exhibiting emerging capabilities such as compositional editing, self-reflective generation, and knowledge-informed generation. These results mark a step toward the next intelligent motion generation. Project Page: https://OmniMoGen.github.io/.
title OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.19159