COIN: Control-Inpainting Diffusion Prior for Human and Camera Motion Estimation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Jiefeng, Yuan, Ye, Rempe, Davis, Zhang, Haotian, Molchanov, Pavlo, Lu, Cewu, Kautz, Jan, Iqbal, Umar
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912006410338304
author Li, Jiefeng
Yuan, Ye
Rempe, Davis
Zhang, Haotian
Molchanov, Pavlo
Lu, Cewu
Kautz, Jan
Iqbal, Umar
author_facet Li, Jiefeng
Yuan, Ye
Rempe, Davis
Zhang, Haotian
Molchanov, Pavlo
Lu, Cewu
Kautz, Jan
Iqbal, Umar
contents Estimating global human motion from moving cameras is challenging due to the entanglement of human and camera motions. To mitigate the ambiguity, existing methods leverage learned human motion priors, which however often result in oversmoothed motions with misaligned 2D projections. To tackle this problem, we propose COIN, a control-inpainting motion diffusion prior that enables fine-grained control to disentangle human and camera motions. Although pre-trained motion diffusion models encode rich motion priors, we find it non-trivial to leverage such knowledge to guide global motion estimation from RGB videos. COIN introduces a novel control-inpainting score distillation sampling method to ensure well-aligned, consistent, and high-quality motion from the diffusion prior within a joint optimization framework. Furthermore, we introduce a new human-scene relation loss to alleviate the scale ambiguity by enforcing consistency among the humans, camera, and scene. Experiments on three challenging benchmarks demonstrate the effectiveness of COIN, which outperforms the state-of-the-art methods in terms of global human motion estimation and camera motion estimation. As an illustrative example, COIN outperforms the state-of-the-art method by 33% in world joint position error (W-MPJPE) on the RICH dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16426
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle COIN: Control-Inpainting Diffusion Prior for Human and Camera Motion Estimation
Li, Jiefeng
Yuan, Ye
Rempe, Davis
Zhang, Haotian
Molchanov, Pavlo
Lu, Cewu
Kautz, Jan
Iqbal, Umar
Computer Vision and Pattern Recognition
Artificial Intelligence
Estimating global human motion from moving cameras is challenging due to the entanglement of human and camera motions. To mitigate the ambiguity, existing methods leverage learned human motion priors, which however often result in oversmoothed motions with misaligned 2D projections. To tackle this problem, we propose COIN, a control-inpainting motion diffusion prior that enables fine-grained control to disentangle human and camera motions. Although pre-trained motion diffusion models encode rich motion priors, we find it non-trivial to leverage such knowledge to guide global motion estimation from RGB videos. COIN introduces a novel control-inpainting score distillation sampling method to ensure well-aligned, consistent, and high-quality motion from the diffusion prior within a joint optimization framework. Furthermore, we introduce a new human-scene relation loss to alleviate the scale ambiguity by enforcing consistency among the humans, camera, and scene. Experiments on three challenging benchmarks demonstrate the effectiveness of COIN, which outperforms the state-of-the-art methods in terms of global human motion estimation and camera motion estimation. As an illustrative example, COIN outperforms the state-of-the-art method by 33% in world joint position error (W-MPJPE) on the RICH dataset.
title COIN: Control-Inpainting Diffusion Prior for Human and Camera Motion Estimation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2408.16426