kNN For Whisper And Its Effect On Bias And Speaker Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nachesa, Maya K., Niculae, Vlad
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915145132802048
author Nachesa, Maya K.
Niculae, Vlad
author_facet Nachesa, Maya K.
Niculae, Vlad
contents Speech recognition performance varies by language, domain, and speaker characteristics such as accent, but fine-tuning a model on any of these categories may lead to catastrophic forgetting. Token-level $k$ nearest neighbor search ($k$NN), first proposed for neural sequence decoders for natural language generation (NLG) and machine translation (MT), is a non-parametric method that instead adapts using inference-time search in an external datastore, without training the underlying model. We show that Whisper, a transformer end-to-end speech model, benefits from $k$NN. We investigate the differences between the speech and text setups. We discuss implications for speaker adaptation, and analyze improvements by gender, accent, and age.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18850
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle kNN For Whisper And Its Effect On Bias And Speaker Adaptation
Nachesa, Maya K.
Niculae, Vlad
Computation and Language
Sound
Audio and Speech Processing
Speech recognition performance varies by language, domain, and speaker characteristics such as accent, but fine-tuning a model on any of these categories may lead to catastrophic forgetting. Token-level $k$ nearest neighbor search ($k$NN), first proposed for neural sequence decoders for natural language generation (NLG) and machine translation (MT), is a non-parametric method that instead adapts using inference-time search in an external datastore, without training the underlying model. We show that Whisper, a transformer end-to-end speech model, benefits from $k$NN. We investigate the differences between the speech and text setups. We discuss implications for speaker adaptation, and analyze improvements by gender, accent, and age.
title kNN For Whisper And Its Effect On Bias And Speaker Adaptation
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.18850