LOSC: LiDAR Open-voc Segmentation Consolidator

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Samet, Nermin, Puy, Gilles, Marlet, Renaud
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910052164567040
author Samet, Nermin
Puy, Gilles
Marlet, Renaud
author_facet Samet, Nermin
Puy, Gilles
Marlet, Renaud
contents We study the use of image-based Vision-Language Models (VLMs) for open-vocabulary segmentation of lidar scans in driving settings. Classically, image semantics can be back-projected onto 3D point clouds. Yet, resulting point labels are noisy and sparse. We consolidate these labels to enforce both spatio-temporal consistency and robustness to image-level augmentations. We then train a 3D network based on these refined labels. This simple method, called LOSC, outperforms the SOTA of zero-shot open-vocabulary semantic and panoptic segmentation on both nuScenes and SemanticKITTI, with significant margins. Code is available at https://github.com/valeoai/LOSC.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LOSC: LiDAR Open-voc Segmentation Consolidator
Samet, Nermin
Puy, Gilles
Marlet, Renaud
Computer Vision and Pattern Recognition
We study the use of image-based Vision-Language Models (VLMs) for open-vocabulary segmentation of lidar scans in driving settings. Classically, image semantics can be back-projected onto 3D point clouds. Yet, resulting point labels are noisy and sparse. We consolidate these labels to enforce both spatio-temporal consistency and robustness to image-level augmentations. We then train a 3D network based on these refined labels. This simple method, called LOSC, outperforms the SOTA of zero-shot open-vocabulary semantic and panoptic segmentation on both nuScenes and SemanticKITTI, with significant margins. Code is available at https://github.com/valeoai/LOSC.
title LOSC: LiDAR Open-voc Segmentation Consolidator
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.07605