kNN For Whisper And Its Effect On Bias And Speaker Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915145132802048 |
|---|---|
| author | Nachesa, Maya K. Niculae, Vlad |
| author_facet | Nachesa, Maya K. Niculae, Vlad |
| contents | Speech recognition performance varies by language, domain, and speaker characteristics such as accent, but fine-tuning a model on any of these categories may lead to catastrophic forgetting. Token-level $k$ nearest neighbor search ($k$NN), first proposed for neural sequence decoders for natural language generation (NLG) and machine translation (MT), is a non-parametric method that instead adapts using inference-time search in an external datastore, without training the underlying model. We show that Whisper, a transformer end-to-end speech model, benefits from $k$NN. We investigate the differences between the speech and text setups. We discuss implications for speaker adaptation, and analyze improvements by gender, accent, and age. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_18850 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | kNN For Whisper And Its Effect On Bias And Speaker Adaptation Nachesa, Maya K. Niculae, Vlad Computation and Language Sound Audio and Speech Processing Speech recognition performance varies by language, domain, and speaker characteristics such as accent, but fine-tuning a model on any of these categories may lead to catastrophic forgetting. Token-level $k$ nearest neighbor search ($k$NN), first proposed for neural sequence decoders for natural language generation (NLG) and machine translation (MT), is a non-parametric method that instead adapts using inference-time search in an external datastore, without training the underlying model. We show that Whisper, a transformer end-to-end speech model, benefits from $k$NN. We investigate the differences between the speech and text setups. We discuss implications for speaker adaptation, and analyze improvements by gender, accent, and age. |
| title | kNN For Whisper And Its Effect On Bias And Speaker Adaptation |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.18850 |