Can Whisper perform speech-based in-context learning?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Siyin, Yang, Chao-Han Huck, Wu, Ji, Zhang, Chao
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910374395117568
author Wang, Siyin
Yang, Chao-Han Huck
Wu, Ji
Zhang, Chao
author_facet Wang, Siyin
Yang, Chao-Han Huck
Wu, Ji
Zhang, Chao
contents This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SICL) approach is proposed for test-time adaptation, which can reduce the word error rates (WERs) with only a small number of labelled speech samples without gradient descent. Language-level adaptation experiments using Chinese dialects showed that when applying SICL to isolated word ASR, consistent and considerable relative WER reductions can be achieved using Whisper models of any size on two dialects, which is on average 32.3%. A k-nearest-neighbours-based in-context example selection technique can be applied to further improve the efficiency of SICL, which can increase the average relative WER reduction to 36.4%. The findings are verified using speaker adaptation or continuous speech recognition tasks, and both achieved considerable relative WER reductions. Detailed quantitative analyses are also provided to shed light on SICL's adaptability to phonological variances and dialect-specific lexical nuances.
format Preprint
id arxiv_https___arxiv_org_abs_2309_07081
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Can Whisper perform speech-based in-context learning?
Wang, Siyin
Yang, Chao-Han Huck
Wu, Ji
Zhang, Chao
Audio and Speech Processing
Computation and Language
Sound
This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SICL) approach is proposed for test-time adaptation, which can reduce the word error rates (WERs) with only a small number of labelled speech samples without gradient descent. Language-level adaptation experiments using Chinese dialects showed that when applying SICL to isolated word ASR, consistent and considerable relative WER reductions can be achieved using Whisper models of any size on two dialects, which is on average 32.3%. A k-nearest-neighbours-based in-context example selection technique can be applied to further improve the efficiency of SICL, which can increase the average relative WER reduction to 36.4%. The findings are verified using speaker adaptation or continuous speech recognition tasks, and both achieved considerable relative WER reductions. Detailed quantitative analyses are also provided to shed light on SICL's adaptability to phonological variances and dialect-specific lexical nuances.
title Can Whisper perform speech-based in-context learning?
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2309.07081