OMUDA: Omni-level Masking for Unsupervised Domain Adaptation in Semantic Segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ou, Yang, Zhao, Xiongwei, Yang, Xinye, Wang, Yihan, Di, Yicheng, Yuan, Rong, Chen, Xieyuanli, Zhu, Xu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909959899316224
author Ou, Yang
Zhao, Xiongwei
Yang, Xinye
Wang, Yihan
Di, Yicheng
Yuan, Rong
Chen, Xieyuanli
Zhu, Xu
author_facet Ou, Yang
Zhao, Xiongwei
Yang, Xinye
Wang, Yihan
Di, Yicheng
Yuan, Rong
Chen, Xieyuanli
Zhu, Xu
contents Unsupervised domain adaptation (UDA) enables semantic segmentation models to generalize from a labeled source domain to an unlabeled target domain. However, existing UDA methods still struggle to bridge the domain gap due to cross-domain contextual ambiguity, inconsistent feature representations, and class-wise pseudo-label noise. To address these challenges, we propose Omni-level Masking for Unsupervised Domain Adaptation (OMUDA), a unified framework that introduces hierarchical masking strategies across distinct representation levels. Specifically, OMUDA comprises: 1) a Context-Aware Masking (CAM) strategy that adaptively distinguishes foreground from background to balance global context and local details; 2) a Feature Distillation Masking (FDM) strategy that enhances robust and consistent feature learning through knowledge transfer from pre-trained models; and 3) a Class Decoupling Masking (CDM) strategy that mitigates the impact of noisy pseudo-labels by explicitly modeling class-wise uncertainty. This hierarchical masking paradigm effectively reduces the domain shift at the contextual, representational, and categorical levels, providing a unified solution beyond existing approaches. Extensive experiments on multiple challenging cross-domain semantic segmentation benchmarks validate the effectiveness of OMUDA. Notably, on the SYNTHIA->Cityscapes and GTA5->Cityscapes tasks, OMUDA can be seamlessly integrated into existing UDA methods and consistently achieving state-of-the-art results with an average improvement of 7%.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12303
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OMUDA: Omni-level Masking for Unsupervised Domain Adaptation in Semantic Segmentation
Ou, Yang
Zhao, Xiongwei
Yang, Xinye
Wang, Yihan
Di, Yicheng
Yuan, Rong
Chen, Xieyuanli
Zhu, Xu
Computer Vision and Pattern Recognition
Unsupervised domain adaptation (UDA) enables semantic segmentation models to generalize from a labeled source domain to an unlabeled target domain. However, existing UDA methods still struggle to bridge the domain gap due to cross-domain contextual ambiguity, inconsistent feature representations, and class-wise pseudo-label noise. To address these challenges, we propose Omni-level Masking for Unsupervised Domain Adaptation (OMUDA), a unified framework that introduces hierarchical masking strategies across distinct representation levels. Specifically, OMUDA comprises: 1) a Context-Aware Masking (CAM) strategy that adaptively distinguishes foreground from background to balance global context and local details; 2) a Feature Distillation Masking (FDM) strategy that enhances robust and consistent feature learning through knowledge transfer from pre-trained models; and 3) a Class Decoupling Masking (CDM) strategy that mitigates the impact of noisy pseudo-labels by explicitly modeling class-wise uncertainty. This hierarchical masking paradigm effectively reduces the domain shift at the contextual, representational, and categorical levels, providing a unified solution beyond existing approaches. Extensive experiments on multiple challenging cross-domain semantic segmentation benchmarks validate the effectiveness of OMUDA. Notably, on the SYNTHIA->Cityscapes and GTA5->Cityscapes tasks, OMUDA can be seamlessly integrated into existing UDA methods and consistently achieving state-of-the-art results with an average improvement of 7%.
title OMUDA: Omni-level Masking for Unsupervised Domain Adaptation in Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.12303