ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bi, Hanbo, Xu, Yulong, Li, Ya, Mao, Yongqiang, Tong, Boyuan, Li, Chongyang, Lang, Chunbo, Diao, Wenhui, Wang, Hongqi, Feng, Yingchao, Sun, Xian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916824209162240
author Bi, Hanbo
Xu, Yulong
Li, Ya
Mao, Yongqiang
Tong, Boyuan
Li, Chongyang
Lang, Chunbo
Diao, Wenhui
Wang, Hongqi
Feng, Yingchao
Sun, Xian
author_facet Bi, Hanbo
Xu, Yulong
Li, Ya
Mao, Yongqiang
Tong, Boyuan
Li, Chongyang
Lang, Chunbo
Diao, Wenhui
Wang, Hongqi
Feng, Yingchao
Sun, Xian
contents The Segment Anything Model (SAM), with its prompt-driven paradigm, exhibits strong generalization in generic segmentation tasks. However, applying SAM to remote sensing (RS) images still faces two major challenges. First, manually constructing precise prompts for each image (e.g., points or boxes) is labor-intensive and inefficient, especially in RS scenarios with dense small objects or spatially fragmented distributions. Second, SAM lacks domain adaptability, as it is pre-trained primarily on natural images and struggles to capture RS-specific semantics and spatial characteristics, especially when segmenting novel or unseen classes. To address these issues, inspired by few-shot learning, we propose ViRefSAM, a novel framework that guides SAM utilizing only a few annotated reference images that contain class-specific objects. Without requiring manual prompts, ViRefSAM enables automatic segmentation of class-consistent objects across RS images. Specifically, ViRefSAM introduces two key components while keeping SAM's original architecture intact: (1) a Visual Contextual Prompt Encoder that extracts class-specific semantic clues from reference images and generates object-aware prompts via contextual interaction with target images; and (2) a Dynamic Target Alignment Adapter, integrated into SAM's image encoder, which mitigates the domain gap by injecting class-specific semantics into target image features, enabling SAM to dynamically focus on task-relevant regions. Extensive experiments on three few-shot segmentation benchmarks, including iSAID-5$^i$, LoveDA-2$^i$, and COCO-20$^i$, demonstrate that ViRefSAM enables accurate and automatic segmentation of unseen classes by leveraging only a few reference images and consistently outperforms existing few-shot segmentation methods across diverse datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02294
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
Bi, Hanbo
Xu, Yulong
Li, Ya
Mao, Yongqiang
Tong, Boyuan
Li, Chongyang
Lang, Chunbo
Diao, Wenhui
Wang, Hongqi
Feng, Yingchao
Sun, Xian
Computer Vision and Pattern Recognition
The Segment Anything Model (SAM), with its prompt-driven paradigm, exhibits strong generalization in generic segmentation tasks. However, applying SAM to remote sensing (RS) images still faces two major challenges. First, manually constructing precise prompts for each image (e.g., points or boxes) is labor-intensive and inefficient, especially in RS scenarios with dense small objects or spatially fragmented distributions. Second, SAM lacks domain adaptability, as it is pre-trained primarily on natural images and struggles to capture RS-specific semantics and spatial characteristics, especially when segmenting novel or unseen classes. To address these issues, inspired by few-shot learning, we propose ViRefSAM, a novel framework that guides SAM utilizing only a few annotated reference images that contain class-specific objects. Without requiring manual prompts, ViRefSAM enables automatic segmentation of class-consistent objects across RS images. Specifically, ViRefSAM introduces two key components while keeping SAM's original architecture intact: (1) a Visual Contextual Prompt Encoder that extracts class-specific semantic clues from reference images and generates object-aware prompts via contextual interaction with target images; and (2) a Dynamic Target Alignment Adapter, integrated into SAM's image encoder, which mitigates the domain gap by injecting class-specific semantics into target image features, enabling SAM to dynamically focus on task-relevant regions. Extensive experiments on three few-shot segmentation benchmarks, including iSAID-5$^i$, LoveDA-2$^i$, and COCO-20$^i$, demonstrate that ViRefSAM enables accurate and automatic segmentation of unseen classes by leveraging only a few reference images and consistently outperforms existing few-shot segmentation methods across diverse datasets.
title ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.02294