KinMo: Kinematic-aware Human Motion Understanding and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Pengfei, Liu, Pinxin, Garrido, Pablo, Kim, Hyeongwoo, Chaudhuri, Bindita
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915422668849152
author Zhang, Pengfei
Liu, Pinxin
Garrido, Pablo
Kim, Hyeongwoo
Chaudhuri, Bindita
author_facet Zhang, Pengfei
Liu, Pinxin
Garrido, Pablo
Kim, Hyeongwoo
Chaudhuri, Bindita
contents Current human motion synthesis frameworks rely on global action descriptions, creating a modality gap that limits both motion understanding and generation capabilities. A single coarse description, such as run, fails to capture details such as variations in speed, limb positioning, and kinematic dynamics, leading to ambiguities between text and motion modalities. To address this challenge, we introduce KinMo, a unified framework built on a hierarchical describable motion representation that extends beyond global actions by incorporating kinematic group movements and their interactions. We design an automated annotation pipeline to generate high-quality, fine-grained descriptions for this decomposition, resulting in the KinMo dataset and offering a scalable and cost-efficient solution for dataset enrichment. To leverage these structured descriptions, we propose Hierarchical Text-Motion Alignment that progressively integrates additional motion details, thereby improving semantic motion understanding. Furthermore, we introduce a coarse-to-fine motion generation procedure to leverage enhanced spatial understanding to improve motion synthesis. Experimental results show that KinMo significantly improves motion understanding, demonstrated by enhanced text-motion retrieval performance and enabling more fine-grained motion generation and editing capabilities. Project Page: https://andypinxinliu.github.io/KinMo
format Preprint
id arxiv_https___arxiv_org_abs_2411_15472
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KinMo: Kinematic-aware Human Motion Understanding and Generation
Zhang, Pengfei
Liu, Pinxin
Garrido, Pablo
Kim, Hyeongwoo
Chaudhuri, Bindita
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Current human motion synthesis frameworks rely on global action descriptions, creating a modality gap that limits both motion understanding and generation capabilities. A single coarse description, such as run, fails to capture details such as variations in speed, limb positioning, and kinematic dynamics, leading to ambiguities between text and motion modalities. To address this challenge, we introduce KinMo, a unified framework built on a hierarchical describable motion representation that extends beyond global actions by incorporating kinematic group movements and their interactions. We design an automated annotation pipeline to generate high-quality, fine-grained descriptions for this decomposition, resulting in the KinMo dataset and offering a scalable and cost-efficient solution for dataset enrichment. To leverage these structured descriptions, we propose Hierarchical Text-Motion Alignment that progressively integrates additional motion details, thereby improving semantic motion understanding. Furthermore, we introduce a coarse-to-fine motion generation procedure to leverage enhanced spatial understanding to improve motion synthesis. Experimental results show that KinMo significantly improves motion understanding, demonstrated by enhanced text-motion retrieval performance and enabling more fine-grained motion generation and editing capabilities. Project Page: https://andypinxinliu.github.io/KinMo
title KinMo: Kinematic-aware Human Motion Understanding and Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
url https://arxiv.org/abs/2411.15472