World-Grounded Human Motion Recovery via Gravity-View Coordinates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Zehong, Pi, Huaijin, Xia, Yan, Cen, Zhi, Peng, Sida, Hu, Zechen, Bao, Hujun, Hu, Ruizhen, Zhou, Xiaowei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929494931013632
author Shen, Zehong
Pi, Huaijin
Xia, Yan
Cen, Zhi
Peng, Sida
Hu, Zechen
Bao, Hujun
Hu, Ruizhen
Zhou, Xiaowei
author_facet Shen, Zehong
Pi, Huaijin
Xia, Yan
Cen, Zhi
Peng, Sida
Hu, Zechen
Bao, Hujun
Hu, Ruizhen
Zhou, Xiaowei
contents We present a novel method for recovering world-grounded human motion from monocular video. The main challenge lies in the ambiguity of defining the world coordinate system, which varies between sequences. Previous approaches attempt to alleviate this issue by predicting relative motion in an autoregressive manner, but are prone to accumulating errors. Instead, we propose estimating human poses in a novel Gravity-View (GV) coordinate system, which is defined by the world gravity and the camera view direction. The proposed GV system is naturally gravity-aligned and uniquely defined for each video frame, largely reducing the ambiguity of learning image-pose mapping. The estimated poses can be transformed back to the world coordinate system using camera rotations, forming a global motion sequence. Additionally, the per-frame estimation avoids error accumulation in the autoregressive methods. Experiments on in-the-wild benchmarks demonstrate that our method recovers more realistic motion in both the camera space and world-grounded settings, outperforming state-of-the-art methods in both accuracy and speed. The code is available at https://zju3dv.github.io/gvhmr/.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06662
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle World-Grounded Human Motion Recovery via Gravity-View Coordinates
Shen, Zehong
Pi, Huaijin
Xia, Yan
Cen, Zhi
Peng, Sida
Hu, Zechen
Bao, Hujun
Hu, Ruizhen
Zhou, Xiaowei
Computer Vision and Pattern Recognition
Artificial Intelligence
We present a novel method for recovering world-grounded human motion from monocular video. The main challenge lies in the ambiguity of defining the world coordinate system, which varies between sequences. Previous approaches attempt to alleviate this issue by predicting relative motion in an autoregressive manner, but are prone to accumulating errors. Instead, we propose estimating human poses in a novel Gravity-View (GV) coordinate system, which is defined by the world gravity and the camera view direction. The proposed GV system is naturally gravity-aligned and uniquely defined for each video frame, largely reducing the ambiguity of learning image-pose mapping. The estimated poses can be transformed back to the world coordinate system using camera rotations, forming a global motion sequence. Additionally, the per-frame estimation avoids error accumulation in the autoregressive methods. Experiments on in-the-wild benchmarks demonstrate that our method recovers more realistic motion in both the camera space and world-grounded settings, outperforming state-of-the-art methods in both accuracy and speed. The code is available at https://zju3dv.github.io/gvhmr/.
title World-Grounded Human Motion Recovery via Gravity-View Coordinates
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2409.06662