ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hwang, Inwoo, Jang, Hojun, Zhou, Bing, Wang, Jian, Kim, Young Min, Guo, Chuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909035370905600
author Hwang, Inwoo
Jang, Hojun
Zhou, Bing
Wang, Jian
Kim, Young Min
Guo, Chuan
author_facet Hwang, Inwoo
Jang, Hojun
Zhou, Bing
Wang, Jian
Kim, Young Min
Guo, Chuan
contents We present ScaleMoGen, a scale-wise autoregressive framework for text-driven human motion generation. Unlike conventional autoregressive approaches that rely on standard next-token prediction, ScaleMoGen frames motion generation as a coarse-to-fine process. We quantize 3D motions into compositional discrete tokens across multiple skeletal-emporal scales of increasing granularity, learning to generate motion by autoregressively predicting next-scale token maps. To maintain structural integrity, our motion tokenizers and quantizers are explicitly designed so that discrete tokens at every scale strictly preserve the skeletal hierarchy. Additionally, we employ bitwise quantization and prediction, which efficiently scale up the tokenizer vocabulary to preserve motion details and stabilize optimization. Extensive experiments demonstrate that ScaleMoGen achieves state-of-the-art performance, establishing an FID of 0.030 (vs. 0.045 for MoMask) on HumanML3D and a CLIP Score of 0.693 (vs. 0.685 for MoMask++) on the SnapMoGen dataset. Furthermore, we demonstrate that our skeletal-temporal multi-scale representation naturally facilitates training-free, text-guided motion editing.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11704
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation
Hwang, Inwoo
Jang, Hojun
Zhou, Bing
Wang, Jian
Kim, Young Min
Guo, Chuan
Computer Vision and Pattern Recognition
We present ScaleMoGen, a scale-wise autoregressive framework for text-driven human motion generation. Unlike conventional autoregressive approaches that rely on standard next-token prediction, ScaleMoGen frames motion generation as a coarse-to-fine process. We quantize 3D motions into compositional discrete tokens across multiple skeletal-emporal scales of increasing granularity, learning to generate motion by autoregressively predicting next-scale token maps. To maintain structural integrity, our motion tokenizers and quantizers are explicitly designed so that discrete tokens at every scale strictly preserve the skeletal hierarchy. Additionally, we employ bitwise quantization and prediction, which efficiently scale up the tokenizer vocabulary to preserve motion details and stabilize optimization. Extensive experiments demonstrate that ScaleMoGen achieves state-of-the-art performance, establishing an FID of 0.030 (vs. 0.045 for MoMask) on HumanML3D and a CLIP Score of 0.693 (vs. 0.685 for MoMask++) on the SnapMoGen dataset. Furthermore, we demonstrate that our skeletal-temporal multi-scale representation naturally facilitates training-free, text-guided motion editing.
title ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.11704