Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915228342550528 |
|---|---|
| author | Liao, Ting-Hsuan Zhou, Yi Shen, Yu Huang, Chun-Hao Paul Mitra, Saayan Huang, Jia-Bin Bhattacharya, Uttaran |
| author_facet | Liao, Ting-Hsuan Zhou, Yi Shen, Yu Huang, Chun-Hao Paul Mitra, Saayan Huang, Jia-Bin Bhattacharya, Uttaran |
| contents | We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and their motion dynamics. Our method addresses this gap by generating body-shape-aware human motions from natural language prompts. We utilize a finite scalar quantization-based variational autoencoder (FSQ-VAE) to quantize motion into discrete tokens and then leverage continuous body shape information to de-quantize these tokens back into continuous, detailed motion. Additionally, we harness the capabilities of a pretrained language model to predict both continuous shape parameters and motion tokens, facilitating the synthesis of text-aligned motions and decoding them into shape-aware motions. We evaluate our method quantitatively and qualitatively, and also conduct a comprehensive perceptual study to demonstrate its efficacy in generating shape-aware motions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_03639 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions Liao, Ting-Hsuan Zhou, Yi Shen, Yu Huang, Chun-Hao Paul Mitra, Saayan Huang, Jia-Bin Bhattacharya, Uttaran Computer Vision and Pattern Recognition We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and their motion dynamics. Our method addresses this gap by generating body-shape-aware human motions from natural language prompts. We utilize a finite scalar quantization-based variational autoencoder (FSQ-VAE) to quantize motion into discrete tokens and then leverage continuous body shape information to de-quantize these tokens back into continuous, detailed motion. Additionally, we harness the capabilities of a pretrained language model to predict both continuous shape parameters and motion tokens, facilitating the synthesis of text-aligned motions and decoding them into shape-aware motions. We evaluate our method quantitatively and qualitatively, and also conduct a comprehensive perceptual study to demonstrate its efficacy in generating shape-aware motions. |
| title | Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.03639 |