The Quest for Generalizable Motion Generation: Data, Model, and Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Jing, Wang, Ruisi, Lu, Junzhe, Huang, Ziqi, Song, Guorui, Zeng, Ailing, Liu, Xian, Wei, Chen, Yin, Wanqi, Sun, Qingping, Cai, Zhongang, Yang, Lei, Liu, Ziwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908918498721792
author Lin, Jing
Wang, Ruisi
Lu, Junzhe
Huang, Ziqi
Song, Guorui
Zeng, Ailing
Liu, Xian
Wei, Chen
Yin, Wanqi
Sun, Qingping
Cai, Zhongang
Yang, Lei
Liu, Ziwei
author_facet Lin, Jing
Wang, Ruisi
Lu, Junzhe
Huang, Ziqi
Song, Guorui
Zeng, Ailing
Liu, Xian
Wei, Chen
Yin, Wanqi
Sun, Qingping
Cai, Zhongang
Yang, Lei
Liu, Ziwei
contents Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most notably video generation (ViGen), have demonstrated remarkable generalization in modeling human behaviors, highlighting transferable insights that MoGen can leverage. Motivated by this observation, we present a comprehensive framework that systematically transfers knowledge from ViGen to MoGen across three key pillars: data, modeling, and evaluation. First, we introduce ViMoGen-228K, a large-scale dataset comprising 228,000 high-quality motion samples that integrates high-fidelity optical MoCap data with semantically annotated motions from web videos and synthesized samples generated by state-of-the-art ViGen models. The dataset includes both text-motion pairs and text-video-motion triplets, substantially expanding semantic diversity. Second, we propose ViMoGen, a flow-matching-based diffusion transformer that unifies priors from MoCap data and ViGen models through gated multimodal conditioning. To enhance efficiency, we further develop ViMoGen-light, a distilled variant that eliminates video generation dependencies while preserving strong generalization. Finally, we present MBench, a hierarchical benchmark designed for fine-grained evaluation across motion quality, prompt fidelity, and generalization ability. Extensive experiments show that our framework significantly outperforms existing approaches in both automatic and human evaluations. The code, data, and benchmark will be made publicly available. Homepage: https://motrixlab.github.io/2026_iclr_vimogen.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
Lin, Jing
Wang, Ruisi
Lu, Junzhe
Huang, Ziqi
Song, Guorui
Zeng, Ailing
Liu, Xian
Wei, Chen
Yin, Wanqi
Sun, Qingping
Cai, Zhongang
Yang, Lei
Liu, Ziwei
Computer Vision and Pattern Recognition
Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most notably video generation (ViGen), have demonstrated remarkable generalization in modeling human behaviors, highlighting transferable insights that MoGen can leverage. Motivated by this observation, we present a comprehensive framework that systematically transfers knowledge from ViGen to MoGen across three key pillars: data, modeling, and evaluation. First, we introduce ViMoGen-228K, a large-scale dataset comprising 228,000 high-quality motion samples that integrates high-fidelity optical MoCap data with semantically annotated motions from web videos and synthesized samples generated by state-of-the-art ViGen models. The dataset includes both text-motion pairs and text-video-motion triplets, substantially expanding semantic diversity. Second, we propose ViMoGen, a flow-matching-based diffusion transformer that unifies priors from MoCap data and ViGen models through gated multimodal conditioning. To enhance efficiency, we further develop ViMoGen-light, a distilled variant that eliminates video generation dependencies while preserving strong generalization. Finally, we present MBench, a hierarchical benchmark designed for fine-grained evaluation across motion quality, prompt fidelity, and generalization ability. Extensive experiments show that our framework significantly outperforms existing approaches in both automatic and human evaluations. The code, data, and benchmark will be made publicly available. Homepage: https://motrixlab.github.io/2026_iclr_vimogen.
title The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.26794