GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Gwanghyun, Li, Xueting, Yuan, Ye, Nagano, Koki, Li, Tianye, Kautz, Jan, Chun, Se Young, Iqbal, Umar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910973324951552
author Kim, Gwanghyun
Li, Xueting
Yuan, Ye
Nagano, Koki
Li, Tianye
Kautz, Jan
Chun, Se Young
Iqbal, Umar
author_facet Kim, Gwanghyun
Li, Xueting
Yuan, Ye
Nagano, Koki
Li, Tianye
Kautz, Jan
Chun, Se Young
Iqbal, Umar
contents Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture fine-grained dynamic details. To address these limitations, we present GeoMan, a novel architecture designed to produce accurate and temporally consistent depth and normal estimations from monocular human videos. GeoMan addresses two key challenges: the scarcity of high-quality 4D training data and the need for metric depth estimation to accurately model human size. To overcome the first challenge, GeoMan employs an image-based model to estimate depth and normals for the first frame of a video, which then conditions a video diffusion model, reframing video geometry estimation task as an image-to-video generation problem. This design offloads the heavy lifting of geometric estimation to the image model and simplifies the video model's role to focus on intricate details while using priors learned from large-scale video datasets. Consequently, GeoMan improves temporal consistency and generalizability while requiring minimal 4D training data. To address the challenge of accurate human size estimation, we introduce a root-relative depth representation that retains critical human-scale details and is easier to be estimated from monocular inputs, overcoming the limitations of traditional affine-invariant and metric depth representations. GeoMan achieves state-of-the-art performance in both qualitative and quantitative evaluations, demonstrating its effectiveness in overcoming longstanding challenges in 3D human geometry estimation from videos.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23085
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion
Kim, Gwanghyun
Li, Xueting
Yuan, Ye
Nagano, Koki
Li, Tianye
Kautz, Jan
Chun, Se Young
Iqbal, Umar
Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture fine-grained dynamic details. To address these limitations, we present GeoMan, a novel architecture designed to produce accurate and temporally consistent depth and normal estimations from monocular human videos. GeoMan addresses two key challenges: the scarcity of high-quality 4D training data and the need for metric depth estimation to accurately model human size. To overcome the first challenge, GeoMan employs an image-based model to estimate depth and normals for the first frame of a video, which then conditions a video diffusion model, reframing video geometry estimation task as an image-to-video generation problem. This design offloads the heavy lifting of geometric estimation to the image model and simplifies the video model's role to focus on intricate details while using priors learned from large-scale video datasets. Consequently, GeoMan improves temporal consistency and generalizability while requiring minimal 4D training data. To address the challenge of accurate human size estimation, we introduce a root-relative depth representation that retains critical human-scale details and is easier to be estimated from monocular inputs, overcoming the limitations of traditional affine-invariant and metric depth representations. GeoMan achieves state-of-the-art performance in both qualitative and quantitative evaluations, demonstrating its effectiveness in overcoming longstanding challenges in 3D human geometry estimation from videos.
title GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
url https://arxiv.org/abs/2505.23085