Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yaoting, Sun, Peiwen, Zhou, Dongzhan, Li, Guangyao, Zhang, Honggang, Hu, Di
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917722350157824
author Wang, Yaoting
Sun, Peiwen
Zhou, Dongzhan
Li, Guangyao
Zhang, Honggang
Hu, Di
author_facet Wang, Yaoting
Sun, Peiwen
Zhou, Dongzhan
Li, Guangyao
Zhang, Honggang
Hu, Di
contents Traditional reference segmentation tasks have predominantly focused on silent visual scenes, neglecting the integral role of multimodal perception and interaction in human experiences. In this work, we introduce a novel task called Reference Audio-Visual Segmentation (Ref-AVS), which seeks to segment objects within the visual domain based on expressions containing multimodal cues. Such expressions are articulated in natural language forms but are enriched with multimodal cues, including audio and visual descriptions. To facilitate this research, we construct the first Ref-AVS benchmark, which provides pixel-level annotations for objects described in corresponding multimodal-cue expressions. To tackle the Ref-AVS task, we propose a new method that adequately utilizes multimodal cues to offer precise segmentation guidance. Finally, we conduct quantitative and qualitative experiments on three test subsets to compare our approach with existing methods from related tasks. The results demonstrate the effectiveness of our method, highlighting its capability to precisely segment objects using multimodal-cue expressions. Dataset is available at \href{https://gewu-lab.github.io/Ref-AVS}{https://gewu-lab.github.io/Ref-AVS}.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10957
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
Wang, Yaoting
Sun, Peiwen
Zhou, Dongzhan
Li, Guangyao
Zhang, Honggang
Hu, Di
Computer Vision and Pattern Recognition
Artificial Intelligence
Traditional reference segmentation tasks have predominantly focused on silent visual scenes, neglecting the integral role of multimodal perception and interaction in human experiences. In this work, we introduce a novel task called Reference Audio-Visual Segmentation (Ref-AVS), which seeks to segment objects within the visual domain based on expressions containing multimodal cues. Such expressions are articulated in natural language forms but are enriched with multimodal cues, including audio and visual descriptions. To facilitate this research, we construct the first Ref-AVS benchmark, which provides pixel-level annotations for objects described in corresponding multimodal-cue expressions. To tackle the Ref-AVS task, we propose a new method that adequately utilizes multimodal cues to offer precise segmentation guidance. Finally, we conduct quantitative and qualitative experiments on three test subsets to compare our approach with existing methods from related tasks. The results demonstrate the effectiveness of our method, highlighting its capability to precisely segment objects using multimodal-cue expressions. Dataset is available at \href{https://gewu-lab.github.io/Ref-AVS}{https://gewu-lab.github.io/Ref-AVS}.
title Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2407.10957