Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910070309126144 |
|---|---|
| author | Abdullah, Badr M. Azime, Israel Abebe Tonja, Atnafu Lambebo Alabi, Jesujoba O. Alemu, Abel Mulat Hagos, Eyob G. Balcha, Bontu Fufa Nerea, Mulubrhan A. Yadeta, Debela Desalegn Marilign, Dagnachew Mekonnen Fentahun, Amanuel Temesgen Kebede, Tadesse Gebru, Israel D. Woldeyohannis, Michael Melese Sewunetie, Walelign Tewabe Möbius, Bernd Klakow, Dietrich |
| author_facet | Abdullah, Badr M. Azime, Israel Abebe Tonja, Atnafu Lambebo Alabi, Jesujoba O. Alemu, Abel Mulat Hagos, Eyob G. Balcha, Bontu Fufa Nerea, Mulubrhan A. Yadeta, Debela Desalegn Marilign, Dagnachew Mekonnen Fentahun, Amanuel Temesgen Kebede, Tadesse Gebru, Israel D. Woldeyohannis, Michael Melese Sewunetie, Walelign Tewabe Möbius, Bernd Klakow, Dietrich |
| contents | We present Ethio-ASR, a suite of multilingual CTC-based automatic speech recognition (ASR) models jointly trained on five Ethiopian languages: Amharic, Tigrinya, Oromo, Sidaama, and Wolaytta. These languages belong to the Semitic, Cushitic, and Omotic branches of the Afroasiatic family, and remain severely underrepresented in speech technology despite being spoken by the vast majority of Ethiopia's population. We train our models on the recently released WAXAL corpus using several pre-trained speech encoders and evaluate against strong multilingual baselines, including OmniASR. Our best model achieves an average WER of 30.48% on the WAXAL test set, outperforming the best OmniASR model with substantially fewer parameters. We further provide a comprehensive analysis of gender bias, the contribution of vowel length and consonant gemination to ASR errors, and the training dynamics of multilingual CTC models. Our models and codebase are publicly available to the research community. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_23654 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages Abdullah, Badr M. Azime, Israel Abebe Tonja, Atnafu Lambebo Alabi, Jesujoba O. Alemu, Abel Mulat Hagos, Eyob G. Balcha, Bontu Fufa Nerea, Mulubrhan A. Yadeta, Debela Desalegn Marilign, Dagnachew Mekonnen Fentahun, Amanuel Temesgen Kebede, Tadesse Gebru, Israel D. Woldeyohannis, Michael Melese Sewunetie, Walelign Tewabe Möbius, Bernd Klakow, Dietrich Computation and Language We present Ethio-ASR, a suite of multilingual CTC-based automatic speech recognition (ASR) models jointly trained on five Ethiopian languages: Amharic, Tigrinya, Oromo, Sidaama, and Wolaytta. These languages belong to the Semitic, Cushitic, and Omotic branches of the Afroasiatic family, and remain severely underrepresented in speech technology despite being spoken by the vast majority of Ethiopia's population. We train our models on the recently released WAXAL corpus using several pre-trained speech encoders and evaluate against strong multilingual baselines, including OmniASR. Our best model achieves an average WER of 30.48% on the WAXAL test set, outperforming the best OmniASR model with substantially fewer parameters. We further provide a comprehensive analysis of gender bias, the contribution of vowel length and consonant gemination to ASR errors, and the training dynamics of multilingual CTC models. Our models and codebase are publicly available to the research community. |
| title | Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2603.23654 |