SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911171101065216 |
|---|---|
| author | Jin, Dian Zhou, Yanghao Zhou, Jinxing Ma, Jiaqi Guo, Ruohao Guo, Dan |
| author_facet | Jin, Dian Zhou, Yanghao Zhou, Jinxing Ma, Jiaqi Guo, Ruohao Guo, Dan |
| contents | Referring Audio-Visual Segmentation (Ref-AVS) aims to segment specific objects in videos based on natural language expressions involving audio, vision, and text information. This task poses significant challenges in cross-modal reasoning and fine-grained object localization. In this paper, we propose a simple framework, SimToken, that integrates a multimodal large language model (MLLM) with the Segment Anything Model (SAM). The MLLM is guided to generate a special semantic token representing the referred object. This compact token, enriched with contextual information from all modalities, acts as a prompt to guide SAM to segment objectsacross video frames. To further improve semantic learning, we introduce a novel target-consistent semantic alignment loss that aligns token embeddings from different expressions but referring to the same object. Experiments on the Ref-AVS benchmark demonstrate that our approach achieves superior performance compared to existing methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_17537 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SimToken: A Simple Baseline for Referring Audio-Visual Segmentation Jin, Dian Zhou, Yanghao Zhou, Jinxing Ma, Jiaqi Guo, Ruohao Guo, Dan Computer Vision and Pattern Recognition Referring Audio-Visual Segmentation (Ref-AVS) aims to segment specific objects in videos based on natural language expressions involving audio, vision, and text information. This task poses significant challenges in cross-modal reasoning and fine-grained object localization. In this paper, we propose a simple framework, SimToken, that integrates a multimodal large language model (MLLM) with the Segment Anything Model (SAM). The MLLM is guided to generate a special semantic token representing the referred object. This compact token, enriched with contextual information from all modalities, acts as a prompt to guide SAM to segment objectsacross video frames. To further improve semantic learning, we introduce a novel target-consistent semantic alignment loss that aligns token embeddings from different expressions but referring to the same object. Experiments on the Ref-AVS benchmark demonstrate that our approach achieves superior performance compared to existing methods. |
| title | SimToken: A Simple Baseline for Referring Audio-Visual Segmentation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.17537 |