Building temporally coherent 3D maps with VGGT for memory-efficient Semantic SLAM

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dinya, Gergely, Halász, Péter, Lőrincz, András, Karacs, Kristóf, Gelencsér-Horváth, Anna
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917107257573376
author Dinya, Gergely
Halász, Péter
Lőrincz, András
Karacs, Kristóf
Gelencsér-Horváth, Anna
author_facet Dinya, Gergely
Halász, Péter
Lőrincz, András
Karacs, Kristóf
Gelencsér-Horváth, Anna
contents We present a fast, spatio-temporal scene understanding framework based on Visual Geometry Grounded Transformer (VGGT). The proposed pipeline is designed to enable efficient, close to real-time performance, supporting applications including assistive navigation. To achieve continuous updates of the 3D scene representation, we process the image flow with a sliding window, aligning submaps, thereby overcoming VGGT's high memory demands. We exploit the VGGT tracking head to aggregate 2D semantic instance masks into 3D objects. To allow for temporal consistency and richer contextual reasoning the system stores timestamps and instance-level identities, thereby enabling the detection of changes in the environment. We evaluate the approach on well-known benchmarks and custom datasets specifically designed for assistive navigation scenarios. The results demonstrate the applicability of the framework to real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16282
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Building temporally coherent 3D maps with VGGT for memory-efficient Semantic SLAM
Dinya, Gergely
Halász, Péter
Lőrincz, András
Karacs, Kristóf
Gelencsér-Horváth, Anna
Computer Vision and Pattern Recognition
We present a fast, spatio-temporal scene understanding framework based on Visual Geometry Grounded Transformer (VGGT). The proposed pipeline is designed to enable efficient, close to real-time performance, supporting applications including assistive navigation. To achieve continuous updates of the 3D scene representation, we process the image flow with a sliding window, aligning submaps, thereby overcoming VGGT's high memory demands. We exploit the VGGT tracking head to aggregate 2D semantic instance masks into 3D objects. To allow for temporal consistency and richer contextual reasoning the system stores timestamps and instance-level identities, thereby enabling the detection of changes in the environment. We evaluate the approach on well-known benchmarks and custom datasets specifically designed for assistive navigation scenarios. The results demonstrate the applicability of the framework to real-world scenarios.
title Building temporally coherent 3D maps with VGGT for memory-efficient Semantic SLAM
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.16282