Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Korse, Srikanth, Elminshawi, Mohamed, Habets, Emanuel A. P., Chetupalli, Srikanth Raj
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911047066058752
author Korse, Srikanth
Elminshawi, Mohamed
Habets, Emanuel A. P.
Chetupalli, Srikanth Raj
author_facet Korse, Srikanth
Elminshawi, Mohamed
Habets, Emanuel A. P.
Chetupalli, Srikanth Raj
contents The primary goal of multi-modal TSE (MTSE) is to extract a target speaker from a speech mixture using complementary information from different modalities, such as audio enrolment and visual feeds corresponding to the target speaker. MTSE systems are expected to perform well even when one of the modalities is unavailable. In practice, the systems often suffer from modality dominance, where one of the modalities outweighs the others, thereby limiting robustness. Our study investigates training strategies and the effect of architectural choices, particularly the normalization layers, in yielding a robust MTSE system in both non-causal and causal configurations. In particular, we propose the use of modality dropout training (MDT) as a superior strategy to standard and multi-task training (MTT) strategies. Experiments conducted on two-speaker mixtures from the LRS3 dataset show the MDT strategy to be effective irrespective of the employed normalization layer. In contrast, the models trained with the standard and MTT strategies are susceptible to modality dominance, and their performance depends on the chosen normalization layer. Additionally, we demonstrate that the system trained with MDT strategy is robust to using extracted speech as the enrollment signal, highlighting its potential applicability in scenarios where the target speaker is not enrolled.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06566
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
Korse, Srikanth
Elminshawi, Mohamed
Habets, Emanuel A. P.
Chetupalli, Srikanth Raj
Audio and Speech Processing
The primary goal of multi-modal TSE (MTSE) is to extract a target speaker from a speech mixture using complementary information from different modalities, such as audio enrolment and visual feeds corresponding to the target speaker. MTSE systems are expected to perform well even when one of the modalities is unavailable. In practice, the systems often suffer from modality dominance, where one of the modalities outweighs the others, thereby limiting robustness. Our study investigates training strategies and the effect of architectural choices, particularly the normalization layers, in yielding a robust MTSE system in both non-causal and causal configurations. In particular, we propose the use of modality dropout training (MDT) as a superior strategy to standard and multi-task training (MTT) strategies. Experiments conducted on two-speaker mixtures from the LRS3 dataset show the MDT strategy to be effective irrespective of the employed normalization layer. In contrast, the models trained with the standard and MTT strategies are susceptible to modality dominance, and their performance depends on the chosen normalization layer. Additionally, we demonstrate that the system trained with MDT strategy is robust to using extracted speech as the enrollment signal, highlighting its potential applicability in scenarios where the target speaker is not enrolled.
title Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
topic Audio and Speech Processing
url https://arxiv.org/abs/2507.06566