Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909624135843840 |
|---|---|
| author | Pan, Zexu Zhao, Shengkui Wang, Tingting Zhou, Kun Ma, Yukun Zhang, Chong Ma, Bin |
| author_facet | Pan, Zexu Zhao, Shengkui Wang, Tingting Zhou, Kun Ma, Yukun Zhang, Chong Ma, Bin |
| contents | Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co-occurring faces are often present on-screen, providing valuable speaker activity cues in the scene. In this work, we introduce a plug-and-play inter-speaker attention module to process these flexible numbers of co-occurring faces, allowing for more accurate speaker extraction in complex multi-person environments. We integrate our module into two prominent models: the AV-DPRNN and the state-of-the-art AV-TFGridNet. Extensive experiments on diverse datasets, including the highly overlapped VoxCeleb2 and sparsely overlapped MISP, demonstrate that our approach consistently outperforms baselines. Furthermore, cross-dataset evaluations on LRS2 and LRS3 confirm the robustness and generalizability of our method. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_20635 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction Pan, Zexu Zhao, Shengkui Wang, Tingting Zhou, Kun Ma, Yukun Zhang, Chong Ma, Bin Audio and Speech Processing Artificial Intelligence Sound Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co-occurring faces are often present on-screen, providing valuable speaker activity cues in the scene. In this work, we introduce a plug-and-play inter-speaker attention module to process these flexible numbers of co-occurring faces, allowing for more accurate speaker extraction in complex multi-person environments. We integrate our module into two prominent models: the AV-DPRNN and the state-of-the-art AV-TFGridNet. Extensive experiments on diverse datasets, including the highly overlapped VoxCeleb2 and sparsely overlapped MISP, demonstrate that our approach consistently outperforms baselines. Furthermore, cross-dataset evaluations on LRS2 and LRS3 confirm the robustness and generalizability of our method. |
| title | Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction |
| topic | Audio and Speech Processing Artificial Intelligence Sound |
| url | https://arxiv.org/abs/2505.20635 |