OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhe, Yuan, Weihao, Shen, Weichao, Zhu, Siyu, Dong, Zilong, Xu, Chang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918161987665920
author Li, Zhe
Yuan, Weihao
Shen, Weichao
Zhu, Siyu
Dong, Zilong
Xu, Chang
author_facet Li, Zhe
Yuan, Weihao
Shen, Weichao
Zhu, Siyu
Dong, Zilong
Xu, Chang
contents Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous methods that usually employ discrete masked modeling or autoregressive modeling, we develop a continuous masked autoregressive motion transformer, where a causal attention is performed considering the sequential nature within the human motion. Within this transformer, we introduce a gated linear attention and an RMSNorm module, which drive the transformer to pay attention to the key actions and suppress the instability caused by either the abnormal movements or the heterogeneous distributions within multi-modalities. To further enhance both the motion generation and the multimodal generalization, we employ the DiT structure to diffuse the conditions from the transformer towards the targets. To fuse different modalities, AdaLN and cross-attention are leveraged to inject the text, speech, and music signals. Experimental results demonstrate that our framework outperforms previous methods across all modalities, including text-to-motion, speech-to-gesture, and music-to-dance. The code of our method will be made public.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14954
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
Li, Zhe
Yuan, Weihao
Shen, Weichao
Zhu, Siyu
Dong, Zilong
Xu, Chang
Computer Vision and Pattern Recognition
Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous methods that usually employ discrete masked modeling or autoregressive modeling, we develop a continuous masked autoregressive motion transformer, where a causal attention is performed considering the sequential nature within the human motion. Within this transformer, we introduce a gated linear attention and an RMSNorm module, which drive the transformer to pay attention to the key actions and suppress the instability caused by either the abnormal movements or the heterogeneous distributions within multi-modalities. To further enhance both the motion generation and the multimodal generalization, we employ the DiT structure to diffuse the conditions from the transformer towards the targets. To fuse different modalities, AdaLN and cross-attention are leveraged to inject the text, speech, and music signals. Experimental results demonstrate that our framework outperforms previous methods across all modalities, including text-to-motion, speech-to-gesture, and music-to-dance. The code of our method will be made public.
title OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.14954