PLANA3R: Zero-shot Metric Planar 3D Reconstruction via Feed-Forward Planar Splatting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Changkun, Tan, Bin, Ke, Zeran, Zhang, Shangzhan, Liu, Jiachen, Qian, Ming, Xue, Nan, Shen, Yujun, Braud, Tristan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918309757190144
author Liu, Changkun
Tan, Bin
Ke, Zeran
Zhang, Shangzhan
Liu, Jiachen
Qian, Ming
Xue, Nan
Shen, Yujun
Braud, Tristan
author_facet Liu, Changkun
Tan, Bin
Ke, Zeran
Zhang, Shangzhan
Liu, Jiachen
Qian, Ming
Xue, Nan
Shen, Yujun
Braud, Tristan
contents This paper addresses metric 3D reconstruction of indoor scenes by exploiting their inherent geometric regularities with compact representations. Using planar 3D primitives - a well-suited representation for man-made environments - we introduce PLANA3R, a pose-free framework for metric Planar 3D Reconstruction from unposed two-view images. Our approach employs Vision Transformers to extract a set of sparse planar primitives, estimate relative camera poses, and supervise geometry learning via planar splatting, where gradients are propagated through high-resolution rendered depth and normal maps of primitives. Unlike prior feedforward methods that require 3D plane annotations during training, PLANA3R learns planar 3D structures without explicit plane supervision, enabling scalable training on large-scale stereo datasets using only depth and normal annotations. We validate PLANA3R on multiple indoor-scene datasets with metric supervision and demonstrate strong generalization to out-of-domain indoor environments across diverse tasks under metric evaluation protocols, including 3D surface reconstruction, depth estimation, and relative pose estimation. Furthermore, by formulating with planar 3D representation, our method emerges with the ability for accurate plane segmentation. The project page is available at https://lck666666.github.io/plana3r
format Preprint
id arxiv_https___arxiv_org_abs_2510_18714
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PLANA3R: Zero-shot Metric Planar 3D Reconstruction via Feed-Forward Planar Splatting
Liu, Changkun
Tan, Bin
Ke, Zeran
Zhang, Shangzhan
Liu, Jiachen
Qian, Ming
Xue, Nan
Shen, Yujun
Braud, Tristan
Computer Vision and Pattern Recognition
This paper addresses metric 3D reconstruction of indoor scenes by exploiting their inherent geometric regularities with compact representations. Using planar 3D primitives - a well-suited representation for man-made environments - we introduce PLANA3R, a pose-free framework for metric Planar 3D Reconstruction from unposed two-view images. Our approach employs Vision Transformers to extract a set of sparse planar primitives, estimate relative camera poses, and supervise geometry learning via planar splatting, where gradients are propagated through high-resolution rendered depth and normal maps of primitives. Unlike prior feedforward methods that require 3D plane annotations during training, PLANA3R learns planar 3D structures without explicit plane supervision, enabling scalable training on large-scale stereo datasets using only depth and normal annotations. We validate PLANA3R on multiple indoor-scene datasets with metric supervision and demonstrate strong generalization to out-of-domain indoor environments across diverse tasks under metric evaluation protocols, including 3D surface reconstruction, depth estimation, and relative pose estimation. Furthermore, by formulating with planar 3D representation, our method emerges with the ability for accurate plane segmentation. The project page is available at https://lck666666.github.io/plana3r
title PLANA3R: Zero-shot Metric Planar 3D Reconstruction via Feed-Forward Planar Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.18714