Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lazarow, Justin, Kang, Kai, Dehghan, Afshin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916766611931136
author Lazarow, Justin
Kang, Kai
Dehghan, Afshin
author_facet Lazarow, Justin
Kang, Kai
Dehghan, Afshin
contents We revisit scene-level 3D object detection as the output of an object-centric framework capable of both localization and mapping using 3D oriented boxes as the underlying geometric primitive. While existing 3D object detection approaches operate globally and implicitly rely on the a priori existence of metric camera poses, our method, Rooms from Motion (RfM) operates on a collection of un-posed images. By replacing the standard 2D keypoint-based matcher of structure-from-motion with an object-centric matcher based on image-derived 3D boxes, we estimate metric camera poses, object tracks, and finally produce a global, semantic 3D object map. When a priori pose is available, we can significantly improve map quality through optimization of global 3D boxes against individual observations. RfM shows strong localization performance and subsequently produces maps of higher quality than leading point-based and multi-view 3D object detection methods on CA-1M and ScanNet++, despite these global methods relying on overparameterization through point clouds or dense volumes. Rooms from Motion achieves a general, object-centric representation which not only extends the work of Cubify Anything to full scenes but also allows for inherently sparse localization and parametric mapping proportional to the number of objects in a scene.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23756
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping
Lazarow, Justin
Kang, Kai
Dehghan, Afshin
Computer Vision and Pattern Recognition
We revisit scene-level 3D object detection as the output of an object-centric framework capable of both localization and mapping using 3D oriented boxes as the underlying geometric primitive. While existing 3D object detection approaches operate globally and implicitly rely on the a priori existence of metric camera poses, our method, Rooms from Motion (RfM) operates on a collection of un-posed images. By replacing the standard 2D keypoint-based matcher of structure-from-motion with an object-centric matcher based on image-derived 3D boxes, we estimate metric camera poses, object tracks, and finally produce a global, semantic 3D object map. When a priori pose is available, we can significantly improve map quality through optimization of global 3D boxes against individual observations. RfM shows strong localization performance and subsequently produces maps of higher quality than leading point-based and multi-view 3D object detection methods on CA-1M and ScanNet++, despite these global methods relying on overparameterization through point clouds or dense volumes. Rooms from Motion achieves a general, object-centric representation which not only extends the work of Cubify Anything to full scenes but also allows for inherently sparse localization and parametric mapping proportional to the number of objects in a scene.
title Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23756