AgMTR: Agent Mining Transformer for Few-shot Segmentation in Remote Sensing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bi, Hanbo, Feng, Yingchao, Mao, Yongqiang, Pei, Jianning, Diao, Wenhui, Wang, Hongqi, Sun, Xian
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914957672579072
author Bi, Hanbo
Feng, Yingchao
Mao, Yongqiang
Pei, Jianning
Diao, Wenhui
Wang, Hongqi
Sun, Xian
author_facet Bi, Hanbo
Feng, Yingchao
Mao, Yongqiang
Pei, Jianning
Diao, Wenhui
Wang, Hongqi
Sun, Xian
contents Few-shot Segmentation (FSS) aims to segment the interested objects in the query image with just a handful of labeled samples (i.e., support images). Previous schemes would leverage the similarity between support-query pixel pairs to construct the pixel-level semantic correlation. However, in remote sensing scenarios with extreme intra-class variations and cluttered backgrounds, such pixel-level correlations may produce tremendous mismatches, resulting in semantic ambiguity between the query foreground (FG) and background (BG) pixels. To tackle this problem, we propose a novel Agent Mining Transformer (AgMTR), which adaptively mines a set of local-aware agents to construct agent-level semantic correlation. Compared with pixel-level semantics, the given agents are equipped with local-contextual information and possess a broader receptive field. At this point, different query pixels can selectively aggregate the fine-grained local semantics of different agents, thereby enhancing the semantic clarity between query FG and BG pixels. Concretely, the Agent Learning Encoder (ALE) is first proposed to erect the optimal transport plan that arranges different agents to aggregate support semantics under different local regions. Then, for further optimizing the agents, the Agent Aggregation Decoder (AAD) and the Semantic Alignment Decoder (SAD) are constructed to break through the limited support set for mining valuable class-specific semantics from unlabeled data sources and the query image itself, respectively. Extensive experiments on the remote sensing benchmark iSAID indicate that the proposed method achieves state-of-the-art performance. Surprisingly, our method remains quite competitive when extended to more common natural scenarios, i.e., PASCAL-5i and COCO-20i.
format Preprint
id arxiv_https___arxiv_org_abs_2409_17453
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AgMTR: Agent Mining Transformer for Few-shot Segmentation in Remote Sensing
Bi, Hanbo
Feng, Yingchao
Mao, Yongqiang
Pei, Jianning
Diao, Wenhui
Wang, Hongqi
Sun, Xian
Computer Vision and Pattern Recognition
Few-shot Segmentation (FSS) aims to segment the interested objects in the query image with just a handful of labeled samples (i.e., support images). Previous schemes would leverage the similarity between support-query pixel pairs to construct the pixel-level semantic correlation. However, in remote sensing scenarios with extreme intra-class variations and cluttered backgrounds, such pixel-level correlations may produce tremendous mismatches, resulting in semantic ambiguity between the query foreground (FG) and background (BG) pixels. To tackle this problem, we propose a novel Agent Mining Transformer (AgMTR), which adaptively mines a set of local-aware agents to construct agent-level semantic correlation. Compared with pixel-level semantics, the given agents are equipped with local-contextual information and possess a broader receptive field. At this point, different query pixels can selectively aggregate the fine-grained local semantics of different agents, thereby enhancing the semantic clarity between query FG and BG pixels. Concretely, the Agent Learning Encoder (ALE) is first proposed to erect the optimal transport plan that arranges different agents to aggregate support semantics under different local regions. Then, for further optimizing the agents, the Agent Aggregation Decoder (AAD) and the Semantic Alignment Decoder (SAD) are constructed to break through the limited support set for mining valuable class-specific semantics from unlabeled data sources and the query image itself, respectively. Extensive experiments on the remote sensing benchmark iSAID indicate that the proposed method achieves state-of-the-art performance. Surprisingly, our method remains quite competitive when extended to more common natural scenarios, i.e., PASCAL-5i and COCO-20i.
title AgMTR: Agent Mining Transformer for Few-shot Segmentation in Remote Sensing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.17453