RemixFusion: Residual-based Mixed Representation for Large-scale Online RGB-D Reconstruction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lan, Yuqing, Zhu, Chenyang, Zhi, Shuaifeng, Zhang, Jiazhao, Wang, Zhoufeng, Yi, Renjiao, Wang, Yijie, Xu, Kai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908539210956800
author Lan, Yuqing
Zhu, Chenyang
Zhi, Shuaifeng
Zhang, Jiazhao
Wang, Zhoufeng
Yi, Renjiao
Wang, Yijie
Xu, Kai
author_facet Lan, Yuqing
Zhu, Chenyang
Zhi, Shuaifeng
Zhang, Jiazhao
Wang, Zhoufeng
Yi, Renjiao
Wang, Yijie
Xu, Kai
contents The introduction of the neural implicit representation has notably propelled the advancement of online dense reconstruction techniques. Compared to traditional explicit representations, such as TSDF, it improves the mapping completeness and memory efficiency. However, the lack of reconstruction details and the time-consuming learning of neural representations hinder the widespread application of neural-based methods to large-scale online reconstruction. We introduce RemixFusion, a novel residual-based mixed representation for scene reconstruction and camera pose estimation dedicated to high-quality and large-scale online RGB-D reconstruction. In particular, we propose a residual-based map representation comprised of an explicit coarse TSDF grid and an implicit neural module that produces residuals representing fine-grained details to be added to the coarse grid. Such mixed representation allows for detail-rich reconstruction with bounded time and memory budget, contrasting with the overly-smoothed results by the purely implicit representations, thus paving the way for high-quality camera tracking. Furthermore, we extend the residual-based representation to handle multi-frame joint pose optimization via bundle adjustment (BA). In contrast to the existing methods, which optimize poses directly, we opt to optimize pose changes. Combined with a novel technique for adaptive gradient amplification, our method attains better optimization convergence and global optimality. Furthermore, we adopt a local moving volume to factorize the mixed scene representation with a divide-and-conquer design to facilitate efficient online learning in our residual-based framework. Extensive experiments demonstrate that our method surpasses all state-of-the-art ones, including those based either on explicit or implicit representations, in terms of the accuracy of both mapping and tracking on large-scale scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RemixFusion: Residual-based Mixed Representation for Large-scale Online RGB-D Reconstruction
Lan, Yuqing
Zhu, Chenyang
Zhi, Shuaifeng
Zhang, Jiazhao
Wang, Zhoufeng
Yi, Renjiao
Wang, Yijie
Xu, Kai
Computer Vision and Pattern Recognition
The introduction of the neural implicit representation has notably propelled the advancement of online dense reconstruction techniques. Compared to traditional explicit representations, such as TSDF, it improves the mapping completeness and memory efficiency. However, the lack of reconstruction details and the time-consuming learning of neural representations hinder the widespread application of neural-based methods to large-scale online reconstruction. We introduce RemixFusion, a novel residual-based mixed representation for scene reconstruction and camera pose estimation dedicated to high-quality and large-scale online RGB-D reconstruction. In particular, we propose a residual-based map representation comprised of an explicit coarse TSDF grid and an implicit neural module that produces residuals representing fine-grained details to be added to the coarse grid. Such mixed representation allows for detail-rich reconstruction with bounded time and memory budget, contrasting with the overly-smoothed results by the purely implicit representations, thus paving the way for high-quality camera tracking. Furthermore, we extend the residual-based representation to handle multi-frame joint pose optimization via bundle adjustment (BA). In contrast to the existing methods, which optimize poses directly, we opt to optimize pose changes. Combined with a novel technique for adaptive gradient amplification, our method attains better optimization convergence and global optimality. Furthermore, we adopt a local moving volume to factorize the mixed scene representation with a divide-and-conquer design to facilitate efficient online learning in our residual-based framework. Extensive experiments demonstrate that our method surpasses all state-of-the-art ones, including those based either on explicit or implicit representations, in terms of the accuracy of both mapping and tracking on large-scale scenes.
title RemixFusion: Residual-based Mixed Representation for Large-scale Online RGB-D Reconstruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.17594