DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Ning, Li, Zhengyu, Loh, Kwong Weng, Xu, Mingxi, Wang, Qi, Wen, Zhengyu, He, Xiaoyu, Zhao, Wei, Gong, Kehong, Zhang, Mingyuan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910013896785920
author Zhang, Ning
Li, Zhengyu
Loh, Kwong Weng
Xu, Mingxi
Wang, Qi
Wen, Zhengyu
He, Xiaoyu
Zhao, Wei
Gong, Kehong
Zhang, Mingyuan
author_facet Zhang, Ning
Li, Zhengyu
Loh, Kwong Weng
Xu, Mingxi
Wang, Qi
Wen, Zhengyu
He, Xiaoyu
Zhao, Wei
Gong, Kehong
Zhang, Mingyuan
contents Prior masked modeling motion generation methods predominantly study text-to-motion. We present DiMo, a discrete diffusion-style framework, which extends masked modeling to bidirectional text--motion understanding and generation. Unlike GPT-style autoregressive approaches that tokenize motion and decode sequentially, DiMo performs iterative masked token refinement, unifying Text-to-Motion (T2M), Motion-to-Text (M2T), and text-free Motion-to-Motion (M2M) within a single model. This decoding paradigm naturally enables a quality-latency trade-off at inference via the number of refinement steps. We further improve motion token fidelity with residual vector quantization (RVQ) and enhance alignment and controllability with Group Relative Policy Optimization (GRPO). Experiments on HumanML3D and KIT-ML show strong motion quality and competitive bidirectional understanding under a unified framework. In addition, we demonstrate model ability in text-free motion completion, text-guided motion prediction and motion caption correction without architectural change. Additional qualitative results are available on our project page: https://animotionlab.github.io/DiMo/.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04188
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding
Zhang, Ning
Li, Zhengyu
Loh, Kwong Weng
Xu, Mingxi
Wang, Qi
Wen, Zhengyu
He, Xiaoyu
Zhao, Wei
Gong, Kehong
Zhang, Mingyuan
Computer Vision and Pattern Recognition
Prior masked modeling motion generation methods predominantly study text-to-motion. We present DiMo, a discrete diffusion-style framework, which extends masked modeling to bidirectional text--motion understanding and generation. Unlike GPT-style autoregressive approaches that tokenize motion and decode sequentially, DiMo performs iterative masked token refinement, unifying Text-to-Motion (T2M), Motion-to-Text (M2T), and text-free Motion-to-Motion (M2M) within a single model. This decoding paradigm naturally enables a quality-latency trade-off at inference via the number of refinement steps. We further improve motion token fidelity with residual vector quantization (RVQ) and enhance alignment and controllability with Group Relative Policy Optimization (GRPO). Experiments on HumanML3D and KIT-ML show strong motion quality and competitive bidirectional understanding under a unified framework. In addition, we demonstrate model ability in text-free motion completion, text-guided motion prediction and motion caption correction without architectural change. Additional qualitative results are available on our project page: https://animotionlab.github.io/DiMo/.
title DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.04188