ParTY: Part-Guidance for Expressive Text-to-Motion Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heo, KunHo, Kim, SuYeon, Gwon, Yonghyun, Kim, Youngbin, Cho, MyeongAh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917329649008640
author Heo, KunHo
Kim, SuYeon
Gwon, Yonghyun
Kim, Youngbin
Cho, MyeongAh
author_facet Heo, KunHo
Kim, SuYeon
Gwon, Yonghyun
Kim, Youngbin
Cho, MyeongAh
contents Text-to-motion synthesis aims to generate natural and expressive human motions from textual descriptions. While existing approaches primarily focus on generating holistic motions from text descriptions, they struggle to accurately reflect actions involving specific body parts. Recent part-wise motion generation methods attempt to resolve this but face two critical limitations: (i) they lack explicit mechanisms for aligning textual semantics with individual body parts, and (ii) they often generate incoherent full-body motions due to integrating independently generated part motions. To overcome these issues and resolve the fundamental trade-off in existing methods, we propose ParTY, a novel framework that enhances part expressiveness while generating coherent full-body motions. ParTY comprises: (1) Part-Guided Network, which first generates part motions to obtain part guidance, then uses it to generate holistic motions; (2) Part-aware Text Grounding, which diversely transforms text embeddings and appropriately aligns them with each body part; and (3) Holistic-Part Fusion, which adaptively fuses holistic motions and part motions. Extensive experiments, including part-level and coherence-level evaluations, demonstrate that ParTY achieves substantial improvements over previous methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09611
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ParTY: Part-Guidance for Expressive Text-to-Motion Synthesis
Heo, KunHo
Kim, SuYeon
Gwon, Yonghyun
Kim, Youngbin
Cho, MyeongAh
Computer Vision and Pattern Recognition
Text-to-motion synthesis aims to generate natural and expressive human motions from textual descriptions. While existing approaches primarily focus on generating holistic motions from text descriptions, they struggle to accurately reflect actions involving specific body parts. Recent part-wise motion generation methods attempt to resolve this but face two critical limitations: (i) they lack explicit mechanisms for aligning textual semantics with individual body parts, and (ii) they often generate incoherent full-body motions due to integrating independently generated part motions. To overcome these issues and resolve the fundamental trade-off in existing methods, we propose ParTY, a novel framework that enhances part expressiveness while generating coherent full-body motions. ParTY comprises: (1) Part-Guided Network, which first generates part motions to obtain part guidance, then uses it to generate holistic motions; (2) Part-aware Text Grounding, which diversely transforms text embeddings and appropriately aligns them with each body part; and (3) Holistic-Part Fusion, which adaptively fuses holistic motions and part motions. Extensive experiments, including part-level and coherence-level evaluations, demonstrate that ParTY achieves substantial improvements over previous methods.
title ParTY: Part-Guidance for Expressive Text-to-Motion Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.09611