Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Weiming, Xiao, Dingwen, Guo, Songyue, Xiang, Guangyu, Wen, Shiqi, Zhao, Minwei, Chen, Lei, Wang, Lin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914461258874880
author Zhang, Weiming
Xiao, Dingwen
Guo, Songyue
Xiang, Guangyu
Wen, Shiqi
Zhao, Minwei
Chen, Lei
Wang, Lin
author_facet Zhang, Weiming
Xiao, Dingwen
Guo, Songyue
Xiang, Guangyu
Wen, Shiqi
Zhao, Minwei
Chen, Lei
Wang, Lin
contents Referring Expression Segmentation (RES) aims to segment image regions described by natural-language expressions, serving as a bridge between vision and language understanding. Existing RES methods, however, rely heavily on large annotated datasets and are limited to either explicit or implicit expressions, hindering their ability to generalize to any referring expression. Recently, the Segment Anything Model 3 (SAM3) has shown impressive robustness in Promptable Concept Segmentation. Nonetheless, applying it to RES remains challenging: (1) SAM3 struggles with longer or implicit expressions; (2) naive coupling of SAM3 with a multimodal large language model (MLLM) makes the final results overly dependent on the MLLM's reasoning capability, without enabling refinement of SAM3's segmentation outputs. To this end, we present Tarot-SAM3, a novel training-free framework that can accurately segment from any referring expression. Specifically, Tarot-SAM3 consists of two key phases. First, the Expression Reasoning Interpreter (ERI) phase introduces reasoning-assisted prompt options to support structured expression parsing and evaluation-aware rephrasing. This transforms arbitrary queries into robust heterogeneous prompts for generating reliable masks with SAM3. Second, the Mask Self-Refining (MSR) phase selects the best mask across prompt types and performs self-refinement by leveraging rich feature relationships from DINOv3 to compare discriminative regions among ERI outputs. It then infers region affiliation to the target, thereby correcting over- and under-segmentation. Extensive experiments demonstrate that Tarot-SAM3 achieves strong performance on both explicit and implicit RES benchmarks, as well as open-world scenarios. Ablation studies further validate the effectiveness of each phase.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07916
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation
Zhang, Weiming
Xiao, Dingwen
Guo, Songyue
Xiang, Guangyu
Wen, Shiqi
Zhao, Minwei
Chen, Lei
Wang, Lin
Computer Vision and Pattern Recognition
Referring Expression Segmentation (RES) aims to segment image regions described by natural-language expressions, serving as a bridge between vision and language understanding. Existing RES methods, however, rely heavily on large annotated datasets and are limited to either explicit or implicit expressions, hindering their ability to generalize to any referring expression. Recently, the Segment Anything Model 3 (SAM3) has shown impressive robustness in Promptable Concept Segmentation. Nonetheless, applying it to RES remains challenging: (1) SAM3 struggles with longer or implicit expressions; (2) naive coupling of SAM3 with a multimodal large language model (MLLM) makes the final results overly dependent on the MLLM's reasoning capability, without enabling refinement of SAM3's segmentation outputs. To this end, we present Tarot-SAM3, a novel training-free framework that can accurately segment from any referring expression. Specifically, Tarot-SAM3 consists of two key phases. First, the Expression Reasoning Interpreter (ERI) phase introduces reasoning-assisted prompt options to support structured expression parsing and evaluation-aware rephrasing. This transforms arbitrary queries into robust heterogeneous prompts for generating reliable masks with SAM3. Second, the Mask Self-Refining (MSR) phase selects the best mask across prompt types and performs self-refinement by leveraging rich feature relationships from DINOv3 to compare discriminative regions among ERI outputs. It then infers region affiliation to the target, thereby correcting over- and under-segmentation. Extensive experiments demonstrate that Tarot-SAM3 achieves strong performance on both explicit and implicit RES benchmarks, as well as open-world scenarios. Ablation studies further validate the effectiveness of each phase.
title Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.07916