InstructSAM: Segment Any Instance with Any Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918533398528000 |
|---|---|
| author | Yuan, Yuqian Li, Wentong Li, Zhaocheng Lin, Yutong Li, Juncheng Tang, Siliang Xiao, Jun Zhuang, Yueting Zhang, Wenqiao |
| author_facet | Yuan, Yuqian Li, Wentong Li, Zhaocheng Lin, Yutong Li, Juncheng Tang, Siliang Xiao, Jun Zhuang, Yueting Zhang, Wenqiao |
| contents | In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions. We formulates instruction-driven instance segmentation as a set-structured query prediction problem and propose an explicit reasoning-to-instance query interface that elegantly bridges a vision-language model (VLM) and SAM3. Specifically, a bank of learnable instance queries is injected into the VLM and contextualized with instruction and visual information, enabling each query to serve as an instance-aware slot. A hybrid-attention mechanism further promotes interaction among these queries, visual tokens, and instruction tokens, improving instance enumeration and reducing duplicate predictions. The resulting LLM-conditioned queries are projected into SAM3's detector query space to drive accurate multi-instance segmentation in a single forward pass. This design equips SAM3 with high-level instruction understanding, compositional reasoning, and instance-level set prediction without modifying its core architecture. To support training and evaluation, we further construct Inst2Seg, a high-quality and large-scale instruction-based instance segmentation dataset and benchmark that couples free-form instructions with instance-level masks. Extensive experiments show that only 2B-scale InstructSAM achieves strong results across complex instruction-driven and phrase-level referring segmentation benchmarks, outperforming prior end-to-end methods and SAM3's agentic pipeline while enabling efficient single-pass multi-instance prediction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_26102 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | InstructSAM: Segment Any Instance with Any Instructions Yuan, Yuqian Li, Wentong Li, Zhaocheng Lin, Yutong Li, Juncheng Tang, Siliang Xiao, Jun Zhuang, Yueting Zhang, Wenqiao Computer Vision and Pattern Recognition In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions. We formulates instruction-driven instance segmentation as a set-structured query prediction problem and propose an explicit reasoning-to-instance query interface that elegantly bridges a vision-language model (VLM) and SAM3. Specifically, a bank of learnable instance queries is injected into the VLM and contextualized with instruction and visual information, enabling each query to serve as an instance-aware slot. A hybrid-attention mechanism further promotes interaction among these queries, visual tokens, and instruction tokens, improving instance enumeration and reducing duplicate predictions. The resulting LLM-conditioned queries are projected into SAM3's detector query space to drive accurate multi-instance segmentation in a single forward pass. This design equips SAM3 with high-level instruction understanding, compositional reasoning, and instance-level set prediction without modifying its core architecture. To support training and evaluation, we further construct Inst2Seg, a high-quality and large-scale instruction-based instance segmentation dataset and benchmark that couples free-form instructions with instance-level masks. Extensive experiments show that only 2B-scale InstructSAM achieves strong results across complex instruction-driven and phrase-level referring segmentation benchmarks, outperforming prior end-to-end methods and SAM3's agentic pipeline while enabling efficient single-pass multi-instance prediction. |
| title | InstructSAM: Segment Any Instance with Any Instructions |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.26102 |