Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Zexu, Zhao, Shengkui, Wang, Tingting, Zhou, Kun, Ma, Yukun, Zhang, Chong, Ma, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909624135843840
author Pan, Zexu
Zhao, Shengkui
Wang, Tingting
Zhou, Kun
Ma, Yukun
Zhang, Chong
Ma, Bin
author_facet Pan, Zexu
Zhao, Shengkui
Wang, Tingting
Zhou, Kun
Ma, Yukun
Zhang, Chong
Ma, Bin
contents Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co-occurring faces are often present on-screen, providing valuable speaker activity cues in the scene. In this work, we introduce a plug-and-play inter-speaker attention module to process these flexible numbers of co-occurring faces, allowing for more accurate speaker extraction in complex multi-person environments. We integrate our module into two prominent models: the AV-DPRNN and the state-of-the-art AV-TFGridNet. Extensive experiments on diverse datasets, including the highly overlapped VoxCeleb2 and sparsely overlapped MISP, demonstrate that our approach consistently outperforms baselines. Furthermore, cross-dataset evaluations on LRS2 and LRS3 confirm the robustness and generalizability of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20635
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
Pan, Zexu
Zhao, Shengkui
Wang, Tingting
Zhou, Kun
Ma, Yukun
Zhang, Chong
Ma, Bin
Audio and Speech Processing
Artificial Intelligence
Sound
Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co-occurring faces are often present on-screen, providing valuable speaker activity cues in the scene. In this work, we introduce a plug-and-play inter-speaker attention module to process these flexible numbers of co-occurring faces, allowing for more accurate speaker extraction in complex multi-person environments. We integrate our module into two prominent models: the AV-DPRNN and the state-of-the-art AV-TFGridNet. Extensive experiments on diverse datasets, including the highly overlapped VoxCeleb2 and sparsely overlapped MISP, demonstrate that our approach consistently outperforms baselines. Furthermore, cross-dataset evaluations on LRS2 and LRS3 confirm the robustness and generalizability of our method.
title Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2505.20635