Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Biaoyu, Wang, Qingsheng, Xu, Cun, Yang, Dingkang, Wang, Wenxuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910252199313408
author Ren, Biaoyu
Wang, Qingsheng
Xu, Cun
Yang, Dingkang
Wang, Wenxuan
author_facet Ren, Biaoyu
Wang, Qingsheng
Xu, Cun
Yang, Dingkang
Wang, Wenxuan
contents Referring Remote Sensing Image Segmentation (RRSIS) is a situated, task-driven cross-modal task related to the embodied perception paradigm, requiring models to align visual-spatial features with linguistic intentions for precise target perception. Recent research has focused on refining the granularity of textual features and optimizing image-text feature fusion to better guide target feature representations. However, insufficient descriptive granularity and sensitivity to semantic shifts can cause bottlenecks in cross-modal feature fusion. To address these issues, we propose the Image-Conditioned Instance Prompt Network (ICIPNet) with Bilateral Information Fusion, which is designed to alleviate bottlenecks in cross-modal feature fusion. ICIPNet introduces an Image-Conditioned Instance Prompt (ICIP) module to generate self-adaptive visual and semantic representations without external knowledge. The Bilateral Information Fusion (BIF) module enhances feature fusion along the token and channel dimensions. Experiments demonstrate that the proposed ICIPNet outperforms existing RRSIS models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24532
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation
Ren, Biaoyu
Wang, Qingsheng
Xu, Cun
Yang, Dingkang
Wang, Wenxuan
Computer Vision and Pattern Recognition
Referring Remote Sensing Image Segmentation (RRSIS) is a situated, task-driven cross-modal task related to the embodied perception paradigm, requiring models to align visual-spatial features with linguistic intentions for precise target perception. Recent research has focused on refining the granularity of textual features and optimizing image-text feature fusion to better guide target feature representations. However, insufficient descriptive granularity and sensitivity to semantic shifts can cause bottlenecks in cross-modal feature fusion. To address these issues, we propose the Image-Conditioned Instance Prompt Network (ICIPNet) with Bilateral Information Fusion, which is designed to alleviate bottlenecks in cross-modal feature fusion. ICIPNet introduces an Image-Conditioned Instance Prompt (ICIP) module to generate self-adaptive visual and semantic representations without external knowledge. The Bilateral Information Fusion (BIF) module enhances feature fusion along the token and channel dimensions. Experiments demonstrate that the proposed ICIPNet outperforms existing RRSIS models.
title Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.24532