End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916448617627648 |
|---|---|
| author | Abdullah, Abdulhady Abas Tabibian, Shima Veisi, Hadi Mahmudi, Aso Rashid, Tarik |
| author_facet | Abdullah, Abdulhady Abas Tabibian, Shima Veisi, Hadi Mahmudi, Aso Rashid, Tarik |
| contents | Automatic Speech Recognition (ASR) for low-resource languages remains a challenging task due to limited training data. This paper introduces a comprehensive study exploring the effectiveness of Whisper, a pre-trained ASR model, for Northern Kurdish (Kurmanji) an under-resourced language spoken in the Middle East. We investigate three fine-tuning strategies: vanilla, specific parameters, and additional modules. Using a Northern Kurdish fine-tuning speech corpus containing approximately 68 hours of validated transcribed data, our experiments demonstrate that the additional module fine-tuning strategy significantly improves ASR accuracy on a specialized test set, achieving a Word Error Rate (WER) of 10.5% and Character Error Rate (CER) of 5.7% with Whisper version 3. These results underscore the potential of sophisticated transformer models for low-resource ASR and emphasize the importance of tailored fine-tuning techniques for optimal performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_16330 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach Abdullah, Abdulhady Abas Tabibian, Shima Veisi, Hadi Mahmudi, Aso Rashid, Tarik Audio and Speech Processing Computation and Language Automatic Speech Recognition (ASR) for low-resource languages remains a challenging task due to limited training data. This paper introduces a comprehensive study exploring the effectiveness of Whisper, a pre-trained ASR model, for Northern Kurdish (Kurmanji) an under-resourced language spoken in the Middle East. We investigate three fine-tuning strategies: vanilla, specific parameters, and additional modules. Using a Northern Kurdish fine-tuning speech corpus containing approximately 68 hours of validated transcribed data, our experiments demonstrate that the additional module fine-tuning strategy significantly improves ASR accuracy on a specialized test set, achieving a Word Error Rate (WER) of 10.5% and Character Error Rate (CER) of 5.7% with Whisper version 3. These results underscore the potential of sophisticated transformer models for low-resource ASR and emphasize the importance of tailored fine-tuning techniques for optimal performance. |
| title | End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach |
| topic | Audio and Speech Processing Computation and Language |
| url | https://arxiv.org/abs/2410.16330 |