InstructSAM: Segment Any Instance with Any Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Yuqian, Li, Wentong, Li, Zhaocheng, Lin, Yutong, Li, Juncheng, Tang, Siliang, Xiao, Jun, Zhuang, Yueting, Zhang, Wenqiao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918533398528000
author Yuan, Yuqian
Li, Wentong
Li, Zhaocheng
Lin, Yutong
Li, Juncheng
Tang, Siliang
Xiao, Jun
Zhuang, Yueting
Zhang, Wenqiao
author_facet Yuan, Yuqian
Li, Wentong
Li, Zhaocheng
Lin, Yutong
Li, Juncheng
Tang, Siliang
Xiao, Jun
Zhuang, Yueting
Zhang, Wenqiao
contents In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions. We formulates instruction-driven instance segmentation as a set-structured query prediction problem and propose an explicit reasoning-to-instance query interface that elegantly bridges a vision-language model (VLM) and SAM3. Specifically, a bank of learnable instance queries is injected into the VLM and contextualized with instruction and visual information, enabling each query to serve as an instance-aware slot. A hybrid-attention mechanism further promotes interaction among these queries, visual tokens, and instruction tokens, improving instance enumeration and reducing duplicate predictions. The resulting LLM-conditioned queries are projected into SAM3's detector query space to drive accurate multi-instance segmentation in a single forward pass. This design equips SAM3 with high-level instruction understanding, compositional reasoning, and instance-level set prediction without modifying its core architecture. To support training and evaluation, we further construct Inst2Seg, a high-quality and large-scale instruction-based instance segmentation dataset and benchmark that couples free-form instructions with instance-level masks. Extensive experiments show that only 2B-scale InstructSAM achieves strong results across complex instruction-driven and phrase-level referring segmentation benchmarks, outperforming prior end-to-end methods and SAM3's agentic pipeline while enabling efficient single-pass multi-instance prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26102
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle InstructSAM: Segment Any Instance with Any Instructions
Yuan, Yuqian
Li, Wentong
Li, Zhaocheng
Lin, Yutong
Li, Juncheng
Tang, Siliang
Xiao, Jun
Zhuang, Yueting
Zhang, Wenqiao
Computer Vision and Pattern Recognition
In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions. We formulates instruction-driven instance segmentation as a set-structured query prediction problem and propose an explicit reasoning-to-instance query interface that elegantly bridges a vision-language model (VLM) and SAM3. Specifically, a bank of learnable instance queries is injected into the VLM and contextualized with instruction and visual information, enabling each query to serve as an instance-aware slot. A hybrid-attention mechanism further promotes interaction among these queries, visual tokens, and instruction tokens, improving instance enumeration and reducing duplicate predictions. The resulting LLM-conditioned queries are projected into SAM3's detector query space to drive accurate multi-instance segmentation in a single forward pass. This design equips SAM3 with high-level instruction understanding, compositional reasoning, and instance-level set prediction without modifying its core architecture. To support training and evaluation, we further construct Inst2Seg, a high-quality and large-scale instruction-based instance segmentation dataset and benchmark that couples free-form instructions with instance-level masks. Extensive experiments show that only 2B-scale InstructSAM achieves strong results across complex instruction-driven and phrase-level referring segmentation benchmarks, outperforming prior end-to-end methods and SAM3's agentic pipeline while enabling efficient single-pass multi-instance prediction.
title InstructSAM: Segment Any Instance with Any Instructions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.26102