Generating Attribute-Aware Human Motions from Textual Prompt

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xinghan, Xu, Kun, Li, Fei, Sheng, Cao, Yu, Jiazhong, Mu, Yadong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908649554706432
author Wang, Xinghan
Xu, Kun
Li, Fei
Sheng, Cao
Yu, Jiazhong
Mu, Yadong
author_facet Wang, Xinghan
Xu, Kun
Li, Fei
Sheng, Cao
Yu, Jiazhong
Mu, Yadong
contents Text-driven human motion generation has recently attracted considerable attention, allowing models to generate human motions based on textual descriptions. However, current methods neglect the influence of human attributes-such as age, gender, weight, and height-which are key factors shaping human motion patterns. This work represents a pilot exploration for bridging this gap. We conceptualize each motion as comprising both attribute information and action semantics, where textual descriptions align exclusively with action semantics. To achieve this, a new framework inspired by Structural Causal Models is proposed to decouple action semantics from human attributes, enabling text-to-semantics prediction and attribute-controlled generation. The resulting model is capable of generating attribute-aware motion aligned with the user's text and attribute inputs. For evaluation, we introduce a comprehensive dataset containing attribute annotations for text-motion pairs, setting the first benchmark for attribute-aware motion generation. Extensive experiments validate our model's effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21912
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Attribute-Aware Human Motions from Textual Prompt
Wang, Xinghan
Xu, Kun
Li, Fei
Sheng, Cao
Yu, Jiazhong
Mu, Yadong
Computer Vision and Pattern Recognition
Multimedia
Text-driven human motion generation has recently attracted considerable attention, allowing models to generate human motions based on textual descriptions. However, current methods neglect the influence of human attributes-such as age, gender, weight, and height-which are key factors shaping human motion patterns. This work represents a pilot exploration for bridging this gap. We conceptualize each motion as comprising both attribute information and action semantics, where textual descriptions align exclusively with action semantics. To achieve this, a new framework inspired by Structural Causal Models is proposed to decouple action semantics from human attributes, enabling text-to-semantics prediction and attribute-controlled generation. The resulting model is capable of generating attribute-aware motion aligned with the user's text and attribute inputs. For evaluation, we introduce a comprehensive dataset containing attribute annotations for text-motion pairs, setting the first benchmark for attribute-aware motion generation. Extensive experiments validate our model's effectiveness.
title Generating Attribute-Aware Human Motions from Textual Prompt
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2506.21912