SimToken: A Simple Baseline for Referring Audio-Visual Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Dian, Zhou, Yanghao, Zhou, Jinxing, Ma, Jiaqi, Guo, Ruohao, Guo, Dan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911171101065216
author Jin, Dian
Zhou, Yanghao
Zhou, Jinxing
Ma, Jiaqi
Guo, Ruohao
Guo, Dan
author_facet Jin, Dian
Zhou, Yanghao
Zhou, Jinxing
Ma, Jiaqi
Guo, Ruohao
Guo, Dan
contents Referring Audio-Visual Segmentation (Ref-AVS) aims to segment specific objects in videos based on natural language expressions involving audio, vision, and text information. This task poses significant challenges in cross-modal reasoning and fine-grained object localization. In this paper, we propose a simple framework, SimToken, that integrates a multimodal large language model (MLLM) with the Segment Anything Model (SAM). The MLLM is guided to generate a special semantic token representing the referred object. This compact token, enriched with contextual information from all modalities, acts as a prompt to guide SAM to segment objectsacross video frames. To further improve semantic learning, we introduce a novel target-consistent semantic alignment loss that aligns token embeddings from different expressions but referring to the same object. Experiments on the Ref-AVS benchmark demonstrate that our approach achieves superior performance compared to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17537
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
Jin, Dian
Zhou, Yanghao
Zhou, Jinxing
Ma, Jiaqi
Guo, Ruohao
Guo, Dan
Computer Vision and Pattern Recognition
Referring Audio-Visual Segmentation (Ref-AVS) aims to segment specific objects in videos based on natural language expressions involving audio, vision, and text information. This task poses significant challenges in cross-modal reasoning and fine-grained object localization. In this paper, we propose a simple framework, SimToken, that integrates a multimodal large language model (MLLM) with the Segment Anything Model (SAM). The MLLM is guided to generate a special semantic token representing the referred object. This compact token, enriched with contextual information from all modalities, acts as a prompt to guide SAM to segment objectsacross video frames. To further improve semantic learning, we introduce a novel target-consistent semantic alignment loss that aligns token embeddings from different expressions but referring to the same object. Experiments on the Ref-AVS benchmark demonstrate that our approach achieves superior performance compared to existing methods.
title SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.17537