Collaborating Foundation Models for Domain Generalized Semantic Segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Benigmim, Yasser, Roy, Subhankar, Essid, Slim, Kalogeiton, Vicky, Lathuilière, Stéphane
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911819880202240
author Benigmim, Yasser
Roy, Subhankar
Essid, Slim
Kalogeiton, Vicky
Lathuilière, Stéphane
author_facet Benigmim, Yasser
Roy, Subhankar
Essid, Slim
Kalogeiton, Vicky
Lathuilière, Stéphane
contents Domain Generalized Semantic Segmentation (DGSS) deals with training a model on a labeled source domain with the aim of generalizing to unseen domains during inference. Existing DGSS methods typically effectuate robust features by means of Domain Randomization (DR). Such an approach is often limited as it can only account for style diversification and not content. In this work, we take an orthogonal approach to DGSS and propose to use an assembly of CoLlaborative FOUndation models for Domain Generalized Semantic Segmentation (CLOUDS). In detail, CLOUDS is a framework that integrates FMs of various kinds: (i) CLIP backbone for its robust feature representation, (ii) generative models to diversify the content, thereby covering various modes of the possible target distribution, and (iii) Segment Anything Model (SAM) for iteratively refining the predictions of the segmentation model. Extensive experiments show that our CLOUDS excels in adapting from synthetic to real DGSS benchmarks and under varying weather conditions, notably outperforming prior methods by 5.6% and 6.7% on averaged miou, respectively. The code is available at : https://github.com/yasserben/CLOUDS
format Preprint
id arxiv_https___arxiv_org_abs_2312_09788
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Collaborating Foundation Models for Domain Generalized Semantic Segmentation
Benigmim, Yasser
Roy, Subhankar
Essid, Slim
Kalogeiton, Vicky
Lathuilière, Stéphane
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Domain Generalized Semantic Segmentation (DGSS) deals with training a model on a labeled source domain with the aim of generalizing to unseen domains during inference. Existing DGSS methods typically effectuate robust features by means of Domain Randomization (DR). Such an approach is often limited as it can only account for style diversification and not content. In this work, we take an orthogonal approach to DGSS and propose to use an assembly of CoLlaborative FOUndation models for Domain Generalized Semantic Segmentation (CLOUDS). In detail, CLOUDS is a framework that integrates FMs of various kinds: (i) CLIP backbone for its robust feature representation, (ii) generative models to diversify the content, thereby covering various modes of the possible target distribution, and (iii) Segment Anything Model (SAM) for iteratively refining the predictions of the segmentation model. Extensive experiments show that our CLOUDS excels in adapting from synthetic to real DGSS benchmarks and under varying weather conditions, notably outperforming prior methods by 5.6% and 6.7% on averaged miou, respectively. The code is available at : https://github.com/yasserben/CLOUDS
title Collaborating Foundation Models for Domain Generalized Semantic Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.09788