Masked Representation Modeling for Domain-Adaptive Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Wenlve, Zhou, Zhiheng, Xian, Tiantao, Zhai, Yikui, Wu, Weibin, Ma, Biyun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918387518537728
author Zhou, Wenlve
Zhou, Zhiheng
Xian, Tiantao
Zhai, Yikui
Wu, Weibin
Ma, Biyun
author_facet Zhou, Wenlve
Zhou, Zhiheng
Xian, Tiantao
Zhai, Yikui
Wu, Weibin
Ma, Biyun
contents Unsupervised domain adaptation (UDA) for semantic segmentation seeks to transfer models from a labeled source domain to an unlabeled target domain. While auxiliary self-supervised tasks such as contrastive learning have enhanced feature discriminability, masked modeling remains underexplored due to architectural constraints and misaligned objectives. We propose Masked Representation Modeling (MRM), an auxiliary task that performs representation masking and reconstruction directly in the latent space. Unlike prior masked modeling methods that reconstruct low-level signals (e.g., pixels or visual tokens), MRM targets high-level semantic features, aligning its objective with segmentation and integrating seamlessly into standard architectures like DeepLab and DAFormer. To support efficient reconstruction, we design a lightweight auxiliary module, Rebuilder, which is jointly trained with the segmentation network but removed during inference, introducing zero test-time overhead. Extensive experiments demonstrate that MRM consistently improves segmentation performance across diverse architectures and UDA benchmarks. When integrated with four representative baselines, MRM achieves an average gain of +2.3 mIoU on GTA $\rightarrow$ Cityscapes and +2.8 mIoU on Cityscapes $\rightarrow$ Synthia, establishing it as a simple, effective, and generalizable strategy for unsupervised domain-adaptive semantic segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13801
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Masked Representation Modeling for Domain-Adaptive Segmentation
Zhou, Wenlve
Zhou, Zhiheng
Xian, Tiantao
Zhai, Yikui
Wu, Weibin
Ma, Biyun
Computer Vision and Pattern Recognition
Unsupervised domain adaptation (UDA) for semantic segmentation seeks to transfer models from a labeled source domain to an unlabeled target domain. While auxiliary self-supervised tasks such as contrastive learning have enhanced feature discriminability, masked modeling remains underexplored due to architectural constraints and misaligned objectives. We propose Masked Representation Modeling (MRM), an auxiliary task that performs representation masking and reconstruction directly in the latent space. Unlike prior masked modeling methods that reconstruct low-level signals (e.g., pixels or visual tokens), MRM targets high-level semantic features, aligning its objective with segmentation and integrating seamlessly into standard architectures like DeepLab and DAFormer. To support efficient reconstruction, we design a lightweight auxiliary module, Rebuilder, which is jointly trained with the segmentation network but removed during inference, introducing zero test-time overhead. Extensive experiments demonstrate that MRM consistently improves segmentation performance across diverse architectures and UDA benchmarks. When integrated with four representative baselines, MRM achieves an average gain of +2.3 mIoU on GTA $\rightarrow$ Cityscapes and +2.8 mIoU on Cityscapes $\rightarrow$ Synthia, establishing it as a simple, effective, and generalizable strategy for unsupervised domain-adaptive semantic segmentation.
title Masked Representation Modeling for Domain-Adaptive Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.13801