Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ying, Kaining, Ding, Henghui, Jie, Guangquan, Jiang, Yu-Gang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915418101252096
author Ying, Kaining
Ding, Henghui
Jie, Guangquan
Jiang, Yu-Gang
author_facet Ying, Kaining
Ding, Henghui
Jie, Guangquan
Jiang, Yu-Gang
contents Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To extend the boundaries of RAVS and facilitate future research in this field, we propose Omnimodal Referring Audio-Visual Segmentation (OmniAVS), a new dataset containing 2,104 videos and 61,095 multimodal referring expressions. OmniAVS stands out with three key innovations: (1) 8 types of multimodal expressions that flexibly combine text, speech, sound, and visual cues; (2) an emphasis on understanding audio content beyond just detecting their presence; and (3) the inclusion of complex reasoning and world knowledge in expressions. Furthermore, we introduce Omnimodal Instructed Segmentation Assistant (OISA), to address the challenges of multimodal reasoning and fine-grained understanding of audiovisual content in OmniAVS. OISA uses MLLM to comprehend complex cues and perform reasoning-based segmentation. Extensive experiments show that OISA outperforms existing methods on OmniAVS and achieves competitive results on other related tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
Ying, Kaining
Ding, Henghui
Jie, Guangquan
Jiang, Yu-Gang
Computer Vision and Pattern Recognition
Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To extend the boundaries of RAVS and facilitate future research in this field, we propose Omnimodal Referring Audio-Visual Segmentation (OmniAVS), a new dataset containing 2,104 videos and 61,095 multimodal referring expressions. OmniAVS stands out with three key innovations: (1) 8 types of multimodal expressions that flexibly combine text, speech, sound, and visual cues; (2) an emphasis on understanding audio content beyond just detecting their presence; and (3) the inclusion of complex reasoning and world knowledge in expressions. Furthermore, we introduce Omnimodal Instructed Segmentation Assistant (OISA), to address the challenges of multimodal reasoning and fine-grained understanding of audiovisual content in OmniAVS. OISA uses MLLM to comprehend complex cues and perform reasoning-based segmentation. Extensive experiments show that OISA outperforms existing methods on OmniAVS and achieves competitive results on other related tasks.
title Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.22886