End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abdullah, Abdulhady Abas, Tabibian, Shima, Veisi, Hadi, Mahmudi, Aso, Rashid, Tarik
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916448617627648
author Abdullah, Abdulhady Abas
Tabibian, Shima
Veisi, Hadi
Mahmudi, Aso
Rashid, Tarik
author_facet Abdullah, Abdulhady Abas
Tabibian, Shima
Veisi, Hadi
Mahmudi, Aso
Rashid, Tarik
contents Automatic Speech Recognition (ASR) for low-resource languages remains a challenging task due to limited training data. This paper introduces a comprehensive study exploring the effectiveness of Whisper, a pre-trained ASR model, for Northern Kurdish (Kurmanji) an under-resourced language spoken in the Middle East. We investigate three fine-tuning strategies: vanilla, specific parameters, and additional modules. Using a Northern Kurdish fine-tuning speech corpus containing approximately 68 hours of validated transcribed data, our experiments demonstrate that the additional module fine-tuning strategy significantly improves ASR accuracy on a specialized test set, achieving a Word Error Rate (WER) of 10.5% and Character Error Rate (CER) of 5.7% with Whisper version 3. These results underscore the potential of sophisticated transformer models for low-resource ASR and emphasize the importance of tailored fine-tuning techniques for optimal performance.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16330
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach
Abdullah, Abdulhady Abas
Tabibian, Shima
Veisi, Hadi
Mahmudi, Aso
Rashid, Tarik
Audio and Speech Processing
Computation and Language
Automatic Speech Recognition (ASR) for low-resource languages remains a challenging task due to limited training data. This paper introduces a comprehensive study exploring the effectiveness of Whisper, a pre-trained ASR model, for Northern Kurdish (Kurmanji) an under-resourced language spoken in the Middle East. We investigate three fine-tuning strategies: vanilla, specific parameters, and additional modules. Using a Northern Kurdish fine-tuning speech corpus containing approximately 68 hours of validated transcribed data, our experiments demonstrate that the additional module fine-tuning strategy significantly improves ASR accuracy on a specialized test set, achieving a Word Error Rate (WER) of 10.5% and Character Error Rate (CER) of 5.7% with Whisper version 3. These results underscore the potential of sophisticated transformer models for low-resource ASR and emphasize the importance of tailored fine-tuning techniques for optimal performance.
title End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2410.16330