Lifting Motion to the 3D World via 2D Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiaman, Liu, C. Karen, Wu, Jiajun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916708306911232
author Li, Jiaman
Liu, C. Karen
Wu, Jiajun
author_facet Li, Jiaman
Liu, C. Karen
Wu, Jiajun
contents Estimating 3D motion from 2D observations is a long-standing research challenge. Prior work typically requires training on datasets containing ground truth 3D motions, limiting their applicability to activities well-represented in existing motion capture data. This dependency particularly hinders generalization to out-of-distribution scenarios or subjects where collecting 3D ground truth is challenging, such as complex athletic movements or animal motion. We introduce MVLift, a novel approach to predict global 3D motion -- including both joint rotations and root trajectories in the world coordinate system -- using only 2D pose sequences for training. Our multi-stage framework leverages 2D motion diffusion models to progressively generate consistent 2D pose sequences across multiple views, a key step in recovering accurate global 3D motion. MVLift generalizes across various domains, including human poses, human-object interactions, and animal poses. Despite not requiring 3D supervision, it outperforms prior work on five datasets, including those methods that require 3D supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18808
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Lifting Motion to the 3D World via 2D Diffusion
Li, Jiaman
Liu, C. Karen
Wu, Jiajun
Computer Vision and Pattern Recognition
Estimating 3D motion from 2D observations is a long-standing research challenge. Prior work typically requires training on datasets containing ground truth 3D motions, limiting their applicability to activities well-represented in existing motion capture data. This dependency particularly hinders generalization to out-of-distribution scenarios or subjects where collecting 3D ground truth is challenging, such as complex athletic movements or animal motion. We introduce MVLift, a novel approach to predict global 3D motion -- including both joint rotations and root trajectories in the world coordinate system -- using only 2D pose sequences for training. Our multi-stage framework leverages 2D motion diffusion models to progressively generate consistent 2D pose sequences across multiple views, a key step in recovering accurate global 3D motion. MVLift generalizes across various domains, including human poses, human-object interactions, and animal poses. Despite not requiring 3D supervision, it outperforms prior work on five datasets, including those methods that require 3D supervision.
title Lifting Motion to the 3D World via 2D Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.18808