SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Zhenjie, Yang, Yuhuan, Ma, Chaofan, Jiang, Dongsheng, Yao, Jiangchao, Zhang, Ya, Wang, Yanfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911288031969280
author Mao, Zhenjie
Yang, Yuhuan
Ma, Chaofan
Jiang, Dongsheng
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
author_facet Mao, Zhenjie
Yang, Yuhuan
Ma, Chaofan
Jiang, Dongsheng
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
contents Referring Image Segmentation (RIS) aims to segment the target object in an image given a natural language expression. While recent methods leverage pre-trained vision backbones and more training corpus to achieve impressive results, they predominantly focus on simple expressions--short, clear noun phrases like "red car" or "left girl". This simplification often reduces RIS to a key word/concept matching problem, limiting the model's ability to handle referential ambiguity in expressions. In this work, we identify two challenging real-world scenarios: object-distracting expressions, which involve multiple entities with contextual cues, and category-implicit expressions, where the object class is not explicitly stated. To address the challenges, we propose a novel framework, SaFiRe, which mimics the human two-phase cognitive process--first forming a global understanding, then refining it through detail-oriented inspection. This is naturally supported by Mamba's scan-then-update property, which aligns with our phased design and enables efficient multi-cycle refinement with linear complexity. We further introduce aRefCOCO, a new benchmark designed to evaluate RIS models under ambiguous referring expressions. Extensive experiments on both standard and proposed datasets demonstrate the superiority of SaFiRe over state-of-the-art baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
Mao, Zhenjie
Yang, Yuhuan
Ma, Chaofan
Jiang, Dongsheng
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Referring Image Segmentation (RIS) aims to segment the target object in an image given a natural language expression. While recent methods leverage pre-trained vision backbones and more training corpus to achieve impressive results, they predominantly focus on simple expressions--short, clear noun phrases like "red car" or "left girl". This simplification often reduces RIS to a key word/concept matching problem, limiting the model's ability to handle referential ambiguity in expressions. In this work, we identify two challenging real-world scenarios: object-distracting expressions, which involve multiple entities with contextual cues, and category-implicit expressions, where the object class is not explicitly stated. To address the challenges, we propose a novel framework, SaFiRe, which mimics the human two-phase cognitive process--first forming a global understanding, then refining it through detail-oriented inspection. This is naturally supported by Mamba's scan-then-update property, which aligns with our phased design and enables efficient multi-cycle refinement with linear complexity. We further introduce aRefCOCO, a new benchmark designed to evaluate RIS models under ambiguous referring expressions. Extensive experiments on both standard and proposed datasets demonstrate the superiority of SaFiRe over state-of-the-art baselines.
title SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.10160