Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mazzucco, Silvio, Persson, Carl, Segu, Mattia, Dovesi, Pier Luigi, Tombari, Federico, Van Gool, Luc, Poggi, Matteo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918150342180864
author Mazzucco, Silvio
Persson, Carl
Segu, Mattia
Dovesi, Pier Luigi
Tombari, Federico
Van Gool, Luc
Poggi, Matteo
author_facet Mazzucco, Silvio
Persson, Carl
Segu, Mattia
Dovesi, Pier Luigi
Tombari, Federico
Van Gool, Luc
Poggi, Matteo
contents We introduce VocAlign, a novel source-free domain adaptation framework specifically designed for VLMs in open-vocabulary semantic segmentation. Our method adopts a student-teacher paradigm enhanced with a vocabulary alignment strategy, which improves pseudo-label generation by incorporating additional class concepts. To ensure efficiency, we use Low-Rank Adaptation (LoRA) to fine-tune the model, preserving its original capabilities while minimizing computational overhead. In addition, we propose a Top-K class selection mechanism for the student model, which significantly reduces memory requirements while further improving adaptation performance. Our approach achieves a notable 6.11 mIoU improvement on the CityScapes dataset and demonstrates superior performance on zero-shot segmentation benchmarks, setting a new standard for source-free adaptation in the open-vocabulary setting.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15225
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
Mazzucco, Silvio
Persson, Carl
Segu, Mattia
Dovesi, Pier Luigi
Tombari, Federico
Van Gool, Luc
Poggi, Matteo
Computer Vision and Pattern Recognition
We introduce VocAlign, a novel source-free domain adaptation framework specifically designed for VLMs in open-vocabulary semantic segmentation. Our method adopts a student-teacher paradigm enhanced with a vocabulary alignment strategy, which improves pseudo-label generation by incorporating additional class concepts. To ensure efficiency, we use Low-Rank Adaptation (LoRA) to fine-tune the model, preserving its original capabilities while minimizing computational overhead. In addition, we propose a Top-K class selection mechanism for the student model, which significantly reduces memory requirements while further improving adaptation performance. Our approach achieves a notable 6.11 mIoU improvement on the CityScapes dataset and demonstrates superior performance on zero-shot segmentation benchmarks, setting a new standard for source-free adaptation in the open-vocabulary setting.
title Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.15225