Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Yue, Yang, Mingyu, Yang, Liuyuxin, Xu, Yang, Yun, Bingxin, Zhang, Yuhe
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910187083792384
author Jiang, Yue
Yang, Mingyu
Yang, Liuyuxin
Xu, Yang
Yun, Bingxin
Zhang, Yuhe
author_facet Jiang, Yue
Yang, Mingyu
Yang, Liuyuxin
Xu, Yang
Yun, Bingxin
Zhang, Yuhe
contents Recent advances in generative motion synthesis have enabled the production of realistic human motions from diverse input modalities. However, synthesizing compound actions from texts, which integrate multiple concurrent actions into coherent full-body sequences, remains a major challenge. We identify two key limitations in current text-to-motion diffusion models: (i) catastrophic neglect, where earlier actions are overwritten by later ones due to improper handling of temporal information, and (ii) attention collapse, which arises from excessive feature fusion in cross-attention mechanisms. As a result, existing approaches often depend on overly detailed textual descriptions (e.g., raising right hand), explicit body-part specifications (e.g., editing the upper body), or the use of large language models (LLMs) for body-part interpretation. These strategies lead to deficient semantic representations of physical structures and kinematic mechanisms, limiting the ability to incorporate natural behaviors such as greeting while walking. To address these issues, we propose the Motion-Adapter, a plug-and-play module that guides text-to-motion diffusion models in generating compound actions by computing decoupled cross-attention maps, which serve as structural masks during the denoising process. Extensive experiments demonstrate that our method consistently produces more faithful and coherent compound motions across diverse textual prompts, surpassing state-of-the-art approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16135
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
Jiang, Yue
Yang, Mingyu
Yang, Liuyuxin
Xu, Yang
Yun, Bingxin
Zhang, Yuhe
Computer Vision and Pattern Recognition
Recent advances in generative motion synthesis have enabled the production of realistic human motions from diverse input modalities. However, synthesizing compound actions from texts, which integrate multiple concurrent actions into coherent full-body sequences, remains a major challenge. We identify two key limitations in current text-to-motion diffusion models: (i) catastrophic neglect, where earlier actions are overwritten by later ones due to improper handling of temporal information, and (ii) attention collapse, which arises from excessive feature fusion in cross-attention mechanisms. As a result, existing approaches often depend on overly detailed textual descriptions (e.g., raising right hand), explicit body-part specifications (e.g., editing the upper body), or the use of large language models (LLMs) for body-part interpretation. These strategies lead to deficient semantic representations of physical structures and kinematic mechanisms, limiting the ability to incorporate natural behaviors such as greeting while walking. To address these issues, we propose the Motion-Adapter, a plug-and-play module that guides text-to-motion diffusion models in generating compound actions by computing decoupled cross-attention maps, which serve as structural masks during the denoising process. Extensive experiments demonstrate that our method consistently produces more faithful and coherent compound motions across diverse textual prompts, surpassing state-of-the-art approaches.
title Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.16135