Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pak, Byeonghyun, Woo, Byeongju, Kim, Sunghwan, Kim, Dae-hwan, Kim, Hoseong
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:https://arxiv.org/abs/2407.09033
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910547942834176
author Pak, Byeonghyun
Woo, Byeongju
Kim, Sunghwan
Kim, Dae-hwan
Kim, Hoseong
author_facet Pak, Byeonghyun
Woo, Byeongju
Kim, Sunghwan
Kim, Dae-hwan
Kim, Hoseong
contents In this paper, we introduce a method to tackle Domain Generalized Semantic Segmentation (DGSS) by utilizing domain-invariant semantic knowledge from text embeddings of vision-language models. We employ the text embeddings as object queries within a transformer-based segmentation framework (textual object queries). These queries are regarded as a domain-invariant basis for pixel grouping in DGSS. To leverage the power of textual object queries, we introduce a novel framework named the textual query-driven mask transformer (tqdm). Our tqdm aims to (1) generate textual object queries that maximally encode domain-invariant semantics and (2) enhance the semantic clarity of dense visual features. Additionally, we suggest three regularization losses to improve the efficacy of tqdm by aligning between visual and textual features. By utilizing our method, the model can comprehend inherent semantic information for classes of interest, enabling it to generalize to extreme domains (e.g., sketch style). Our tqdm achieves 68.9 mIoU on GTA5$\rightarrow$Cityscapes, outperforming the prior state-of-the-art method by 2.5 mIoU. The project page is available at https://byeonghyunpak.github.io/tqdm.
format Preprint
id arxiv_https___arxiv_org_abs_2407_09033
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
Pak, Byeonghyun
Woo, Byeongju
Kim, Sunghwan
Kim, Dae-hwan
Kim, Hoseong
Computer Vision and Pattern Recognition
In this paper, we introduce a method to tackle Domain Generalized Semantic Segmentation (DGSS) by utilizing domain-invariant semantic knowledge from text embeddings of vision-language models. We employ the text embeddings as object queries within a transformer-based segmentation framework (textual object queries). These queries are regarded as a domain-invariant basis for pixel grouping in DGSS. To leverage the power of textual object queries, we introduce a novel framework named the textual query-driven mask transformer (tqdm). Our tqdm aims to (1) generate textual object queries that maximally encode domain-invariant semantics and (2) enhance the semantic clarity of dense visual features. Additionally, we suggest three regularization losses to improve the efficacy of tqdm by aligning between visual and textual features. By utilizing our method, the model can comprehend inherent semantic information for classes of interest, enabling it to generalize to extreme domains (e.g., sketch style). Our tqdm achieves 68.9 mIoU on GTA5$\rightarrow$Cityscapes, outperforming the prior state-of-the-art method by 2.5 mIoU. The project page is available at https://byeonghyunpak.github.io/tqdm.
title Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.09033