Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abu, Turi, Shi, Ying, Zheng, Thomas Fang, Wang, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909471979077632
author Abu, Turi
Shi, Ying
Zheng, Thomas Fang
Wang, Dong
author_facet Abu, Turi
Shi, Ying
Zheng, Thomas Fang
Wang, Dong
contents We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowd-sourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours of real-world audio recordings paired with transcriptions, covering read speech in both clean and noisy environments. This dataset addresses the critical need for ASR resources for the Oromo language which is underrepresented. To show its applicability for the ASR task, we conducted experiments using the Conformer model, achieving a Word Error Rate (WER) of 15.32% with hybrid CTC and AED loss and WER of 18.74% with pure CTC loss. Additionally, fine-tuning the Whisper model resulted in a significantly improved WER of 10.82%. These results establish baselines for Oromo ASR, highlighting both the challenges and the potential for improving ASR performance in Oromo. The dataset is publicly available at https://github.com/turinaf/sagalee and we encourage its use for further research and development in Oromo speech processing.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00421
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
Abu, Turi
Shi, Ying
Zheng, Thomas Fang
Wang, Dong
Computation and Language
Sound
Audio and Speech Processing
We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowd-sourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours of real-world audio recordings paired with transcriptions, covering read speech in both clean and noisy environments. This dataset addresses the critical need for ASR resources for the Oromo language which is underrepresented. To show its applicability for the ASR task, we conducted experiments using the Conformer model, achieving a Word Error Rate (WER) of 15.32% with hybrid CTC and AED loss and WER of 18.74% with pure CTC loss. Additionally, fine-tuning the Whisper model resulted in a significantly improved WER of 10.82%. These results establish baselines for Oromo ASR, highlighting both the challenges and the potential for improving ASR performance in Oromo. The dataset is publicly available at https://github.com/turinaf/sagalee and we encourage its use for further research and development in Oromo speech processing.
title Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2502.00421