Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Linyi, Tucker, Richard, Li, Zhengqi, Fouhey, David, Snavely, Noah, Holynski, Aleksander
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912354196783104
author Jin, Linyi
Tucker, Richard
Li, Zhengqi
Fouhey, David
Snavely, Noah
Holynski, Aleksander
author_facet Jin, Linyi
Tucker, Richard
Li, Zhengqi
Fouhey, David
Snavely, Noah
Holynski, Aleksander
contents Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly supervising methods for recovering 3D motion remains challenging due to the fundamental difficulty of obtaining ground truth annotations. We present a system for mining high-quality 4D reconstructions from internet stereoscopic, wide-angle videos. Our system fuses and filters the outputs of camera pose estimation, stereo depth estimation, and temporal tracking methods into high-quality dynamic 3D reconstructions. We use this method to generate large-scale data in the form of world-consistent, pseudo-metric 3D point clouds with long-term motion trajectories. We demonstrate the utility of this data by training a variant of DUSt3R to predict structure and 3D motion from real-world image pairs, showing that training on our reconstructed data enables generalization to diverse real-world scenes. Project page and data at: https://stereo4d.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2412_09621
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
Jin, Linyi
Tucker, Richard
Li, Zhengqi
Fouhey, David
Snavely, Noah
Holynski, Aleksander
Computer Vision and Pattern Recognition
Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly supervising methods for recovering 3D motion remains challenging due to the fundamental difficulty of obtaining ground truth annotations. We present a system for mining high-quality 4D reconstructions from internet stereoscopic, wide-angle videos. Our system fuses and filters the outputs of camera pose estimation, stereo depth estimation, and temporal tracking methods into high-quality dynamic 3D reconstructions. We use this method to generate large-scale data in the form of world-consistent, pseudo-metric 3D point clouds with long-term motion trajectories. We demonstrate the utility of this data by training a variant of DUSt3R to predict structure and 3D motion from real-world image pairs, showing that training on our reconstructed data enables generalization to diverse real-world scenes. Project page and data at: https://stereo4d.github.io
title Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09621