Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vesterbacka, Leonora, Rekathati, Faton, Kurtz, Robin, Sikora, Justyna, Toftgård, Agnes
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909736370176000
author Vesterbacka, Leonora
Rekathati, Faton
Kurtz, Robin
Sikora, Justyna
Toftgård, Agnes
author_facet Vesterbacka, Leonora
Rekathati, Faton
Kurtz, Robin
Sikora, Justyna
Toftgård, Agnes
contents This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual training datasets, substantial improvements in performance can be achieved by fine-tuning existing multilingual models, as shown in this work. This work reports an overall improvement across model sizes compared to OpenAI's Whisper evaluated on Swedish. Most notably, we report an average 47% reduction in WER comparing our best performing model to OpenAI's whisper-large-v3, in evaluations across FLEURS, Common Voice, and NST.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17538
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
Vesterbacka, Leonora
Rekathati, Faton
Kurtz, Robin
Sikora, Justyna
Toftgård, Agnes
Computation and Language
Sound
Audio and Speech Processing
This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual training datasets, substantial improvements in performance can be achieved by fine-tuning existing multilingual models, as shown in this work. This work reports an overall improvement across model sizes compared to OpenAI's Whisper evaluated on Swedish. Most notably, we report an average 47% reduction in WER comparing our best performing model to OpenAI's whisper-large-v3, in evaluations across FLEURS, Common Voice, and NST.
title Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.17538