SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Qi, Wei, Dongxu, Zhao, Lingzhe, Li, Wenpu, Huang, Zhangchi, Ji, Shunping, Liu, Peidong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908559960178688
author Xu, Qi
Wei, Dongxu
Zhao, Lingzhe
Li, Wenpu
Huang, Zhangchi
Ji, Shunping
Liu, Peidong
author_facet Xu, Qi
Wei, Dongxu
Zhao, Lingzhe
Li, Wenpu
Huang, Zhangchi
Ji, Shunping
Liu, Peidong
contents Simultaneous understanding and 3D reconstruction plays an important role in developing end-to-end embodied intelligent systems. To achieve this, recent approaches resort to 2D-to-3D feature alignment paradigm, which leads to limited 3D understanding capability and potential semantic information loss. In light of this, we propose SIU3R, the first alignment-free framework for generalizable simultaneous understanding and 3D reconstruction from unposed images. Specifically, SIU3R bridges reconstruction and understanding tasks via pixel-aligned 3D representation, and unifies multiple understanding (segmentation) tasks into a set of unified learnable queries, enabling native 3D understanding without the need of alignment with 2D models. To encourage collaboration between the two tasks with shared representation, we further conduct in-depth analyses of their mutual benefits, and propose two lightweight modules to facilitate their interaction. Extensive experiments demonstrate that our method achieves state-of-the-art performance not only on the individual tasks of 3D reconstruction and understanding, but also on the task of simultaneous understanding and 3D reconstruction, highlighting the advantages of our alignment-free framework and the effectiveness of the mutual benefit designs. Project page: https://insomniaaac.github.io/siu3r/
format Preprint
id arxiv_https___arxiv_org_abs_2507_02705
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature Alignment
Xu, Qi
Wei, Dongxu
Zhao, Lingzhe
Li, Wenpu
Huang, Zhangchi
Ji, Shunping
Liu, Peidong
Computer Vision and Pattern Recognition
Simultaneous understanding and 3D reconstruction plays an important role in developing end-to-end embodied intelligent systems. To achieve this, recent approaches resort to 2D-to-3D feature alignment paradigm, which leads to limited 3D understanding capability and potential semantic information loss. In light of this, we propose SIU3R, the first alignment-free framework for generalizable simultaneous understanding and 3D reconstruction from unposed images. Specifically, SIU3R bridges reconstruction and understanding tasks via pixel-aligned 3D representation, and unifies multiple understanding (segmentation) tasks into a set of unified learnable queries, enabling native 3D understanding without the need of alignment with 2D models. To encourage collaboration between the two tasks with shared representation, we further conduct in-depth analyses of their mutual benefits, and propose two lightweight modules to facilitate their interaction. Extensive experiments demonstrate that our method achieves state-of-the-art performance not only on the individual tasks of 3D reconstruction and understanding, but also on the task of simultaneous understanding and 3D reconstruction, highlighting the advantages of our alignment-free framework and the effectiveness of the mutual benefit designs. Project page: https://insomniaaac.github.io/siu3r/
title SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.02705