RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duisterhof, Bardienus P., Oberst, Jan, Wen, Bowen, Birchfield, Stan, Ramanan, Deva, Ichnowski, Jeffrey
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916781230129152
author Duisterhof, Bardienus P.
Oberst, Jan
Wen, Bowen
Birchfield, Stan
Ramanan, Deva
Ichnowski, Jeffrey
author_facet Duisterhof, Bardienus P.
Oberst, Jan
Wen, Bowen
Birchfield, Stan
Ramanan, Deva
Ichnowski, Jeffrey
contents 3D shape completion has broad applications in robotics, digital twin reconstruction, and extended reality (XR). Although recent advances in 3D object and scene completion have achieved impressive results, existing methods lack 3D consistency, are computationally expensive, and struggle to capture sharp object boundaries. Our work (RaySt3R) addresses these limitations by recasting 3D shape completion as a novel view synthesis problem. Specifically, given a single RGB-D image and a novel viewpoint (encoded as a collection of query rays), we train a feedforward transformer to predict depth maps, object masks, and per-pixel confidence scores for those query rays. RaySt3R fuses these predictions across multiple query views to reconstruct complete 3D shapes. We evaluate RaySt3R on synthetic and real-world datasets, and observe it achieves state-of-the-art performance, outperforming the baselines on all datasets by up to 44% in 3D chamfer distance. Project page: https://rayst3r.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2506_05285
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion
Duisterhof, Bardienus P.
Oberst, Jan
Wen, Bowen
Birchfield, Stan
Ramanan, Deva
Ichnowski, Jeffrey
Computer Vision and Pattern Recognition
3D shape completion has broad applications in robotics, digital twin reconstruction, and extended reality (XR). Although recent advances in 3D object and scene completion have achieved impressive results, existing methods lack 3D consistency, are computationally expensive, and struggle to capture sharp object boundaries. Our work (RaySt3R) addresses these limitations by recasting 3D shape completion as a novel view synthesis problem. Specifically, given a single RGB-D image and a novel viewpoint (encoded as a collection of query rays), we train a feedforward transformer to predict depth maps, object masks, and per-pixel confidence scores for those query rays. RaySt3R fuses these predictions across multiple query views to reconstruct complete 3D shapes. We evaluate RaySt3R on synthetic and real-world datasets, and observe it achieves state-of-the-art performance, outperforming the baselines on all datasets by up to 44% in 3D chamfer distance. Project page: https://rayst3r.github.io
title RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05285