Dynamic Visual SLAM using a General 3D Prior

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhong, Xingguang, Jin, Liren, Popović, Marija, Behley, Jens, Stachniss, Cyrill
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918236205875200
author Zhong, Xingguang
Jin, Liren
Popović, Marija
Behley, Jens
Stachniss, Cyrill
author_facet Zhong, Xingguang
Jin, Liren
Popović, Marija
Behley, Jens
Stachniss, Cyrill
contents Reliable incremental estimation of camera poses and 3D reconstruction is key to enable various applications including robotics, interactive visualization, and augmented reality. However, this task is particularly challenging in dynamic natural environments, where scene dynamics can severely deteriorate camera pose estimation accuracy. In this work, we propose a novel monocular visual SLAM system that can robustly estimate camera poses in dynamic scenes. To this end, we leverage the complementary strengths of geometric patch-based online bundle adjustment and recent feed-forward reconstruction models. Specifically, we propose a feed-forward reconstruction model to precisely filter out dynamic regions, while also utilizing its depth prediction to enhance the robustness of the patch-based visual SLAM. By aligning depth prediction with estimated patches from bundle adjustment, we robustly handle the inherent scale ambiguities of the batch-wise application of the feed-forward reconstruction model.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06868
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Visual SLAM using a General 3D Prior
Zhong, Xingguang
Jin, Liren
Popović, Marija
Behley, Jens
Stachniss, Cyrill
Robotics
Computer Vision and Pattern Recognition
Reliable incremental estimation of camera poses and 3D reconstruction is key to enable various applications including robotics, interactive visualization, and augmented reality. However, this task is particularly challenging in dynamic natural environments, where scene dynamics can severely deteriorate camera pose estimation accuracy. In this work, we propose a novel monocular visual SLAM system that can robustly estimate camera poses in dynamic scenes. To this end, we leverage the complementary strengths of geometric patch-based online bundle adjustment and recent feed-forward reconstruction models. Specifically, we propose a feed-forward reconstruction model to precisely filter out dynamic regions, while also utilizing its depth prediction to enhance the robustness of the patch-based visual SLAM. By aligning depth prediction with estimated patches from bundle adjustment, we robustly handle the inherent scale ambiguities of the batch-wise application of the feed-forward reconstruction model.
title Dynamic Visual SLAM using a General 3D Prior
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.06868