CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mata, Cristina, Ranasinghe, Kanchana, Ryoo, Michael S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911048289746944
author Mata, Cristina
Ranasinghe, Kanchana
Ryoo, Michael S.
author_facet Mata, Cristina
Ranasinghe, Kanchana
Ryoo, Michael S.
contents Unsupervised domain adaptation (UDA) involves learning class semantics from labeled data within a source domain that generalize to an unseen target domain. UDA methods are particularly impactful for semantic segmentation, where annotations are more difficult to collect than in image classification. Despite recent advances in large-scale vision-language representation learning, UDA methods for segmentation have not taken advantage of the domain-agnostic properties of text. To address this, we present a novel Covariance-based Pixel-Text loss, CoPT, that uses domain-agnostic text embeddings to learn domain-invariant features in an image segmentation encoder. The text embeddings are generated through our LLM Domain Template process, where an LLM is used to generate source and target domain descriptions that are fed to a frozen CLIP model and combined. In experiments on four benchmarks we show that a model trained using CoPT achieves the new state of the art performance on UDA for segmentation. The code can be found at https://github.com/cfmata/CoPT.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07125
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings
Mata, Cristina
Ranasinghe, Kanchana
Ryoo, Michael S.
Computer Vision and Pattern Recognition
Image and Video Processing
Unsupervised domain adaptation (UDA) involves learning class semantics from labeled data within a source domain that generalize to an unseen target domain. UDA methods are particularly impactful for semantic segmentation, where annotations are more difficult to collect than in image classification. Despite recent advances in large-scale vision-language representation learning, UDA methods for segmentation have not taken advantage of the domain-agnostic properties of text. To address this, we present a novel Covariance-based Pixel-Text loss, CoPT, that uses domain-agnostic text embeddings to learn domain-invariant features in an image segmentation encoder. The text embeddings are generated through our LLM Domain Template process, where an LLM is used to generate source and target domain descriptions that are fed to a frozen CLIP model and combined. In experiments on four benchmarks we show that a model trained using CoPT achieves the new state of the art performance on UDA for segmentation. The code can be found at https://github.com/cfmata/CoPT.
title CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2507.07125