Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Zhan, Zeng, Bang, Yang, Peijun, Du, Jiarong, Ju, Wei, Tian, Yao, Liu, Juan, Li, Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912959703285760
author Jin, Zhan
Zeng, Bang
Yang, Peijun
Du, Jiarong
Ju, Wei
Tian, Yao
Liu, Juan
Li, Ming
author_facet Jin, Zhan
Zeng, Bang
Yang, Peijun
Du, Jiarong
Ju, Wei
Tian, Yao
Liu, Juan
Li, Ming
contents Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images, and frame-level lip motion or facial expression features --can significantly improve performance. However, real-world applications often suffer from intermittent signal loss, especially for frame-level cues. This paper systematically investigates the robustness of multi-enrollment fusion under varying degrees of modality missing. Results show that while full multimodal fusion excels under ideal conditions, its performance degrades sharply when encountering unseen modalities missing during the testing. Crucially, training with a high missing rate dramatically enhances robustness, maintaining stable performance even under severe test-time modality missing. We demonstrate that fusing the complementary one frame of face image with frame-level lip features achieves both strong performance and robustness for the AVTSE task. The model and codes are shared.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12583
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
Jin, Zhan
Zeng, Bang
Yang, Peijun
Du, Jiarong
Ju, Wei
Tian, Yao
Liu, Juan
Li, Ming
Audio and Speech Processing
Sound
Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images, and frame-level lip motion or facial expression features --can significantly improve performance. However, real-world applications often suffer from intermittent signal loss, especially for frame-level cues. This paper systematically investigates the robustness of multi-enrollment fusion under varying degrees of modality missing. Results show that while full multimodal fusion excels under ideal conditions, its performance degrades sharply when encountering unseen modalities missing during the testing. Crucially, training with a high missing rate dramatically enhances robustness, maintaining stable performance even under severe test-time modality missing. We demonstrate that fusing the complementary one frame of face image with frame-level lip features achieves both strong performance and robustness for the AVTSE task. The model and codes are shared.
title Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.12583