Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Jinxing, Zhou, Yanghao, Han, Mingfei, Wang, Tong, Chang, Xiaojun, Cholakkal, Hisham, Anwer, Rao Muhammad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908480206536704
author Zhou, Jinxing
Zhou, Yanghao
Han, Mingfei
Wang, Tong
Chang, Xiaojun
Cholakkal, Hisham
Anwer, Rao Muhammad
author_facet Zhou, Jinxing
Zhou, Yanghao
Han, Mingfei
Wang, Tong
Chang, Xiaojun
Cholakkal, Hisham
Anwer, Rao Muhammad
contents Referring Audio-Visual Segmentation (Ref-AVS) aims to segment target objects in audible videos based on given reference expressions. Prior works typically rely on learning latent embeddings via multimodal fusion to prompt a tunable SAM/SAM2 decoder for segmentation, which requires strong pixel-level supervision and lacks interpretability. From a novel perspective of explicit reference understanding, we propose TGS-Agent, which decomposes the task into a Think-Ground-Segment process, mimicking the human reasoning procedure by first identifying the referred object through multimodal analysis, followed by coarse-grained grounding and precise segmentation. To this end, we first propose Ref-Thinker, a multimodal language model capable of reasoning over textual, visual, and auditory cues. We construct an instruction-tuning dataset with explicit object-aware think-answer chains for Ref-Thinker fine-tuning. The object description inferred by Ref-Thinker is used as an explicit prompt for Grounding-DINO and SAM2, which perform grounding and segmentation without relying on pixel-level supervision. Additionally, we introduce R\textsuperscript{2}-AVSBench, a new benchmark with linguistically diverse and reasoning-intensive references for better evaluating model generalization. Our approach achieves state-of-the-art results on both standard Ref-AVSBench and proposed R\textsuperscript{2}-AVSBench. Code will be available at https://github.com/jasongief/TGS-Agent.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04418
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
Zhou, Jinxing
Zhou, Yanghao
Han, Mingfei
Wang, Tong
Chang, Xiaojun
Cholakkal, Hisham
Anwer, Rao Muhammad
Multimedia
Computer Vision and Pattern Recognition
Multiagent Systems
Sound
Audio and Speech Processing
Referring Audio-Visual Segmentation (Ref-AVS) aims to segment target objects in audible videos based on given reference expressions. Prior works typically rely on learning latent embeddings via multimodal fusion to prompt a tunable SAM/SAM2 decoder for segmentation, which requires strong pixel-level supervision and lacks interpretability. From a novel perspective of explicit reference understanding, we propose TGS-Agent, which decomposes the task into a Think-Ground-Segment process, mimicking the human reasoning procedure by first identifying the referred object through multimodal analysis, followed by coarse-grained grounding and precise segmentation. To this end, we first propose Ref-Thinker, a multimodal language model capable of reasoning over textual, visual, and auditory cues. We construct an instruction-tuning dataset with explicit object-aware think-answer chains for Ref-Thinker fine-tuning. The object description inferred by Ref-Thinker is used as an explicit prompt for Grounding-DINO and SAM2, which perform grounding and segmentation without relying on pixel-level supervision. Additionally, we introduce R\textsuperscript{2}-AVSBench, a new benchmark with linguistically diverse and reasoning-intensive references for better evaluating model generalization. Our approach achieves state-of-the-art results on both standard Ref-AVSBench and proposed R\textsuperscript{2}-AVSBench. Code will be available at https://github.com/jasongief/TGS-Agent.
title Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
topic Multimedia
Computer Vision and Pattern Recognition
Multiagent Systems
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.04418