Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Alavilli, Sagarika, Banerjee, Annesya, Elbanna, Gasser, Magaro, Annika
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910638702329856
author Alavilli, Sagarika
Banerjee, Annesya
Elbanna, Gasser
Magaro, Annika
author_facet Alavilli, Sagarika
Banerjee, Annesya
Elbanna, Gasser
Magaro, Annika
contents Current state-of-the-art speech recognition models are trained to map acoustic signals into sub-lexical units. While these models demonstrate superior performance, they remain vulnerable to out-of-distribution conditions such as background noise and speech augmentations. In this work, we hypothesize that incorporating speaker representations during speech recognition can enhance model robustness to noise. We developed a transformer-based model that jointly performs speech recognition and speaker identification. Our model utilizes speech embeddings from Whisper and speaker embeddings from ECAPA-TDNN, which are processed jointly to perform both tasks. We show that the joint model performs comparably to Whisper under clean conditions. Notably, the joint model outperforms Whisper in high-noise environments, such as with 8-speaker babble background noise. Furthermore, our joint model excels in handling highly augmented speech, including sine-wave and noise-vocoded speech. Overall, these results suggest that integrating voice representations with speech recognition can lead to more robust models under adversarial conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05423
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
Alavilli, Sagarika
Banerjee, Annesya
Elbanna, Gasser
Magaro, Annika
Sound
Artificial Intelligence
Audio and Speech Processing
Current state-of-the-art speech recognition models are trained to map acoustic signals into sub-lexical units. While these models demonstrate superior performance, they remain vulnerable to out-of-distribution conditions such as background noise and speech augmentations. In this work, we hypothesize that incorporating speaker representations during speech recognition can enhance model robustness to noise. We developed a transformer-based model that jointly performs speech recognition and speaker identification. Our model utilizes speech embeddings from Whisper and speaker embeddings from ECAPA-TDNN, which are processed jointly to perform both tasks. We show that the joint model performs comparably to Whisper under clean conditions. Notably, the joint model outperforms Whisper in high-noise environments, such as with 8-speaker babble background noise. Furthermore, our joint model excels in handling highly augmented speech, including sine-wave and noise-vocoded speech. Overall, these results suggest that integrating voice representations with speech recognition can lead to more robust models under adversarial conditions.
title Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2410.05423