Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liao, Ting-Hsuan, Zhou, Yi, Shen, Yu, Huang, Chun-Hao Paul, Mitra, Saayan, Huang, Jia-Bin, Bhattacharya, Uttaran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915228342550528
author Liao, Ting-Hsuan
Zhou, Yi
Shen, Yu
Huang, Chun-Hao Paul
Mitra, Saayan
Huang, Jia-Bin
Bhattacharya, Uttaran
author_facet Liao, Ting-Hsuan
Zhou, Yi
Shen, Yu
Huang, Chun-Hao Paul
Mitra, Saayan
Huang, Jia-Bin
Bhattacharya, Uttaran
contents We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and their motion dynamics. Our method addresses this gap by generating body-shape-aware human motions from natural language prompts. We utilize a finite scalar quantization-based variational autoencoder (FSQ-VAE) to quantize motion into discrete tokens and then leverage continuous body shape information to de-quantize these tokens back into continuous, detailed motion. Additionally, we harness the capabilities of a pretrained language model to predict both continuous shape parameters and motion tokens, facilitating the synthesis of text-aligned motions and decoding them into shape-aware motions. We evaluate our method quantitatively and qualitatively, and also conduct a comprehensive perceptual study to demonstrate its efficacy in generating shape-aware motions.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions
Liao, Ting-Hsuan
Zhou, Yi
Shen, Yu
Huang, Chun-Hao Paul
Mitra, Saayan
Huang, Jia-Bin
Bhattacharya, Uttaran
Computer Vision and Pattern Recognition
We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and their motion dynamics. Our method addresses this gap by generating body-shape-aware human motions from natural language prompts. We utilize a finite scalar quantization-based variational autoencoder (FSQ-VAE) to quantize motion into discrete tokens and then leverage continuous body shape information to de-quantize these tokens back into continuous, detailed motion. Additionally, we harness the capabilities of a pretrained language model to predict both continuous shape parameters and motion tokens, facilitating the synthesis of text-aligned motions and decoding them into shape-aware motions. We evaluate our method quantitatively and qualitatively, and also conduct a comprehensive perceptual study to demonstrate its efficacy in generating shape-aware motions.
title Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.03639