MinkUNeXt: Point Cloud-based Large-scale Place Recognition using 3D Sparse Convolutions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cabrera, J. J., Santo, A., Gil, A., Viegas, C., Payá, L.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909363115917312
author Cabrera, J. J.
Santo, A.
Gil, A.
Viegas, C.
Payá, L.
author_facet Cabrera, J. J.
Santo, A.
Gil, A.
Viegas, C.
Payá, L.
contents This paper presents MinkUNeXt, an effective and efficient architecture for place-recognition from point clouds entirely based on the new 3D MinkNeXt Block, a residual block composed of 3D sparse convolutions that follows the philosophy established by recent Transformers but purely using simple 3D convolutions. Feature extraction is performed at different scales by a U-Net encoder-decoder network and the feature aggregation of those features into a single descriptor is carried out by a Generalized Mean Pooling (GeM). The proposed architecture demonstrates that it is possible to surpass the current state-of-the-art by only relying on conventional 3D sparse convolutions without making use of more complex and sophisticated proposals such as Transformers, Attention-Layers or Deformable Convolutions. A thorough assessment of the proposal has been carried out using the Oxford RobotCar and the In-house datasets. As a result, MinkUNeXt proves to outperform other methods in the state-of-the-art.
format Preprint
id arxiv_https___arxiv_org_abs_2403_07593
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MinkUNeXt: Point Cloud-based Large-scale Place Recognition using 3D Sparse Convolutions
Cabrera, J. J.
Santo, A.
Gil, A.
Viegas, C.
Payá, L.
Computer Vision and Pattern Recognition
This paper presents MinkUNeXt, an effective and efficient architecture for place-recognition from point clouds entirely based on the new 3D MinkNeXt Block, a residual block composed of 3D sparse convolutions that follows the philosophy established by recent Transformers but purely using simple 3D convolutions. Feature extraction is performed at different scales by a U-Net encoder-decoder network and the feature aggregation of those features into a single descriptor is carried out by a Generalized Mean Pooling (GeM). The proposed architecture demonstrates that it is possible to surpass the current state-of-the-art by only relying on conventional 3D sparse convolutions without making use of more complex and sophisticated proposals such as Transformers, Attention-Layers or Deformable Convolutions. A thorough assessment of the proposal has been carried out using the Oxford RobotCar and the In-house datasets. As a result, MinkUNeXt proves to outperform other methods in the state-of-the-art.
title MinkUNeXt: Point Cloud-based Large-scale Place Recognition using 3D Sparse Convolutions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.07593