TMT: Cross-domain Semantic Segmentation with Region-adaptive Transferability Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Enming, Li, Zhengyu, Wu, Yanru, Wang, Jingge, Tan, Yang, Wang, Guan, Li, Yang, Zhang, Xiaoping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917014384148480
author Zhang, Enming
Li, Zhengyu
Wu, Yanru
Wang, Jingge
Tan, Yang
Wang, Guan
Li, Yang
Zhang, Xiaoping
author_facet Zhang, Enming
Li, Zhengyu
Wu, Yanru
Wang, Jingge
Tan, Yang
Wang, Guan
Li, Yang
Zhang, Xiaoping
contents Recent advances in Vision Transformers (ViTs) have significantly advanced semantic segmentation performance. However, their adaptation to new target domains remains challenged by distribution shifts, which often disrupt global attention mechanisms. While existing global and patch-level adaptation methods offer some improvements, they overlook the spatially varying transferability inherent in different image regions. To address this, we propose the Transferable Mask Transformer (TMT), a region-adaptive framework designed to enhance cross-domain representation learning through transferability guidance. First, we dynamically partition the image into coherent regions, grouped by structural and semantic similarity, and estimates their domain transferability at a localized level. Then, we incorporate region-level transferability maps directly into the self-attention mechanism of ViTs, allowing the model to adaptively focus attention on areas with lower transferability and higher semantic uncertainty. Extensive experiments across 20 diverse cross-domain settings demonstrate that TMT not only mitigates the performance degradation typically associated with domain shift but also consistently outperforms existing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05774
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TMT: Cross-domain Semantic Segmentation with Region-adaptive Transferability Estimation
Zhang, Enming
Li, Zhengyu
Wu, Yanru
Wang, Jingge
Tan, Yang
Wang, Guan
Li, Yang
Zhang, Xiaoping
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in Vision Transformers (ViTs) have significantly advanced semantic segmentation performance. However, their adaptation to new target domains remains challenged by distribution shifts, which often disrupt global attention mechanisms. While existing global and patch-level adaptation methods offer some improvements, they overlook the spatially varying transferability inherent in different image regions. To address this, we propose the Transferable Mask Transformer (TMT), a region-adaptive framework designed to enhance cross-domain representation learning through transferability guidance. First, we dynamically partition the image into coherent regions, grouped by structural and semantic similarity, and estimates their domain transferability at a localized level. Then, we incorporate region-level transferability maps directly into the self-attention mechanism of ViTs, allowing the model to adaptively focus attention on areas with lower transferability and higher semantic uncertainty. Extensive experiments across 20 diverse cross-domain settings demonstrate that TMT not only mitigates the performance degradation typically associated with domain shift but also consistently outperforms existing approaches.
title TMT: Cross-domain Semantic Segmentation with Region-adaptive Transferability Estimation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.05774