IGLOSS: Image Generation for Lidar Open-vocabulary Semantic Segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Samet, Nermin, Puy, Gilles, Marlet, Renaud
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914439176912896
author Samet, Nermin
Puy, Gilles
Marlet, Renaud
author_facet Samet, Nermin
Puy, Gilles
Marlet, Renaud
contents This paper presents a new method for the zero-shot open-vocabulary semantic segmentation (OVSS) of 3D automotive lidar data. To circumvent the recognized image-text modality gap that is intrinsic to approaches based on Vision Language Models (VLMs) such as CLIP, our method relies instead on image generation from text, to create prototype images. Given a 3D network distilled from a 2D Vision Foundation Model (VFM), we then label a point cloud by matching 3D point features with 2D image features of these prototypes. Our method is state-of-the-art for OVSS on nuScenes and SemanticKITTI. Code, pre-trained models, and generated images are available at https://github.com/valeoai/IGLOSS.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01361
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IGLOSS: Image Generation for Lidar Open-vocabulary Semantic Segmentation
Samet, Nermin
Puy, Gilles
Marlet, Renaud
Computer Vision and Pattern Recognition
This paper presents a new method for the zero-shot open-vocabulary semantic segmentation (OVSS) of 3D automotive lidar data. To circumvent the recognized image-text modality gap that is intrinsic to approaches based on Vision Language Models (VLMs) such as CLIP, our method relies instead on image generation from text, to create prototype images. Given a 3D network distilled from a 2D Vision Foundation Model (VFM), we then label a point cloud by matching 3D point features with 2D image features of these prototypes. Our method is state-of-the-art for OVSS on nuScenes and SemanticKITTI. Code, pre-trained models, and generated images are available at https://github.com/valeoai/IGLOSS.
title IGLOSS: Image Generation for Lidar Open-vocabulary Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.01361