ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ganlin, Qian, Shenhan, Wang, Xi, Cremers, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911356782903296
author Zhang, Ganlin
Qian, Shenhan
Wang, Xi
Cremers, Daniel
author_facet Zhang, Ganlin
Qian, Shenhan
Wang, Xi
Cremers, Daniel
contents We present ViSTA-SLAM as a real-time monocular visual SLAM system that operates without requiring camera intrinsics, making it broadly applicable across diverse camera setups. At its core, the system employs a lightweight symmetric two-view association (STA) model as the frontend, which simultaneously estimates relative camera poses and regresses local pointmaps from only two RGB images. This design reduces model complexity significantly, the size of our frontend is only 35\% that of comparable state-of-the-art methods, while enhancing the quality of two-view constraints used in the pipeline. In the backend, we construct a specially designed Sim(3) pose graph that incorporates loop closures to address accumulated drift. Extensive experiments demonstrate that our approach achieves superior performance in both camera tracking and dense 3D reconstruction quality compared to current methods. Github repository: https://github.com/zhangganlin/vista-slam
format Preprint
id arxiv_https___arxiv_org_abs_2509_01584
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association
Zhang, Ganlin
Qian, Shenhan
Wang, Xi
Cremers, Daniel
Computer Vision and Pattern Recognition
We present ViSTA-SLAM as a real-time monocular visual SLAM system that operates without requiring camera intrinsics, making it broadly applicable across diverse camera setups. At its core, the system employs a lightweight symmetric two-view association (STA) model as the frontend, which simultaneously estimates relative camera poses and regresses local pointmaps from only two RGB images. This design reduces model complexity significantly, the size of our frontend is only 35\% that of comparable state-of-the-art methods, while enhancing the quality of two-view constraints used in the pipeline. In the backend, we construct a specially designed Sim(3) pose graph that incorporates loop closures to address accumulated drift. Extensive experiments demonstrate that our approach achieves superior performance in both camera tracking and dense 3D reconstruction quality compared to current methods. Github repository: https://github.com/zhangganlin/vista-slam
title ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.01584