ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Guanghao, Ren, Kerui, Xu, Linning, Zheng, Zhewen, Jiang, Changjian, Gao, Xin, Dai, Bo, Pu, Jian, Yu, Mulin, Pang, Jiangmiao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915542651109376
author Li, Guanghao
Ren, Kerui
Xu, Linning
Zheng, Zhewen
Jiang, Changjian
Gao, Xin
Dai, Bo
Pu, Jian
Yu, Mulin
Pang, Jiangmiao
author_facet Li, Guanghao
Ren, Kerui
Xu, Linning
Zheng, Zhewen
Jiang, Changjian
Gao, Xin
Dai, Bo
Pu, Jian
Yu, Mulin
Pang, Jiangmiao
contents On-the-fly 3D reconstruction from monocular image sequences is a long-standing challenge in computer vision, critical for applications such as real-to-sim, AR/VR, and robotics. Existing methods face a major tradeoff: per-scene optimization yields high fidelity but is computationally expensive, whereas feed-forward foundation models enable real-time inference but struggle with accuracy and robustness. In this work, we propose ARTDECO, a unified framework that combines the efficiency of feed-forward models with the reliability of SLAM-based pipelines. ARTDECO uses 3D foundation models for pose estimation and point prediction, coupled with a Gaussian decoder that transforms multi-scale features into structured 3D Gaussians. To sustain both fidelity and efficiency at scale, we design a hierarchical Gaussian representation with a LoD-aware rendering strategy, which improves rendering fidelity while reducing redundancy. Experiments on eight diverse indoor and outdoor benchmarks show that ARTDECO delivers interactive performance comparable to SLAM, robustness similar to feed-forward systems, and reconstruction quality close to per-scene optimization, providing a practical path toward on-the-fly digitization of real-world environments with both accurate geometry and high visual fidelity. Explore more demos on our project page: https://city-super.github.io/artdeco/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08551
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation
Li, Guanghao
Ren, Kerui
Xu, Linning
Zheng, Zhewen
Jiang, Changjian
Gao, Xin
Dai, Bo
Pu, Jian
Yu, Mulin
Pang, Jiangmiao
Computer Vision and Pattern Recognition
On-the-fly 3D reconstruction from monocular image sequences is a long-standing challenge in computer vision, critical for applications such as real-to-sim, AR/VR, and robotics. Existing methods face a major tradeoff: per-scene optimization yields high fidelity but is computationally expensive, whereas feed-forward foundation models enable real-time inference but struggle with accuracy and robustness. In this work, we propose ARTDECO, a unified framework that combines the efficiency of feed-forward models with the reliability of SLAM-based pipelines. ARTDECO uses 3D foundation models for pose estimation and point prediction, coupled with a Gaussian decoder that transforms multi-scale features into structured 3D Gaussians. To sustain both fidelity and efficiency at scale, we design a hierarchical Gaussian representation with a LoD-aware rendering strategy, which improves rendering fidelity while reducing redundancy. Experiments on eight diverse indoor and outdoor benchmarks show that ARTDECO delivers interactive performance comparable to SLAM, robustness similar to feed-forward systems, and reconstruction quality close to per-scene optimization, providing a practical path toward on-the-fly digitization of real-world environments with both accurate geometry and high visual fidelity. Explore more demos on our project page: https://city-super.github.io/artdeco/.
title ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.08551