HS-SLAM: Hybrid Representation with Structural Supervision for Improved Dense SLAM

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gong, Ziren, Tosi, Fabio, Zhang, Youmin, Mattoccia, Stefano, Poggi, Matteo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912297705799680
author Gong, Ziren
Tosi, Fabio
Zhang, Youmin
Mattoccia, Stefano
Poggi, Matteo
author_facet Gong, Ziren
Tosi, Fabio
Zhang, Youmin
Mattoccia, Stefano
Poggi, Matteo
contents NeRF-based SLAM has recently achieved promising results in tracking and reconstruction. However, existing methods face challenges in providing sufficient scene representation, capturing structural information, and maintaining global consistency in scenes emerging significant movement or being forgotten. To this end, we present HS-SLAM to tackle these problems. To enhance scene representation capacity, we propose a hybrid encoding network that combines the complementary strengths of hash-grid, tri-planes, and one-blob, improving the completeness and smoothness of reconstruction. Additionally, we introduce structural supervision by sampling patches of non-local pixels rather than individual rays to better capture the scene structure. To ensure global consistency, we implement an active global bundle adjustment (BA) to eliminate camera drifts and mitigate accumulative errors. Experimental results demonstrate that HS-SLAM outperforms the baselines in tracking and reconstruction accuracy while maintaining the efficiency required for robotics.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21778
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HS-SLAM: Hybrid Representation with Structural Supervision for Improved Dense SLAM
Gong, Ziren
Tosi, Fabio
Zhang, Youmin
Mattoccia, Stefano
Poggi, Matteo
Computer Vision and Pattern Recognition
NeRF-based SLAM has recently achieved promising results in tracking and reconstruction. However, existing methods face challenges in providing sufficient scene representation, capturing structural information, and maintaining global consistency in scenes emerging significant movement or being forgotten. To this end, we present HS-SLAM to tackle these problems. To enhance scene representation capacity, we propose a hybrid encoding network that combines the complementary strengths of hash-grid, tri-planes, and one-blob, improving the completeness and smoothness of reconstruction. Additionally, we introduce structural supervision by sampling patches of non-local pixels rather than individual rays to better capture the scene structure. To ensure global consistency, we implement an active global bundle adjustment (BA) to eliminate camera drifts and mitigate accumulative errors. Experimental results demonstrate that HS-SLAM outperforms the baselines in tracking and reconstruction accuracy while maintaining the efficiency required for robotics.
title HS-SLAM: Hybrid Representation with Structural Supervision for Improved Dense SLAM
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.21778