The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paraskevopoulos, Georgios, Tsoukala, Chara, Katsamanis, Athanasios, Katsouros, Vassilis
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917700882661376
author Paraskevopoulos, Georgios
Tsoukala, Chara
Katsamanis, Athanasios
Katsouros, Vassilis
author_facet Paraskevopoulos, Georgios
Tsoukala, Chara
Katsamanis, Athanasios
Katsouros, Vassilis
contents The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue is exacerbated in the era of large, data-intensive models. Recent research has underscored the potential of leveraging weak supervision to augment the pool of available data. In this study, we compile an 800-hour corpus of Modern Greek from podcasts and employ Whisper large-v3 to generate silver transcriptions. This corpus is utilized to fine-tune our models, aiming to assess the efficacy of this approach in enhancing ASR performance. Our analysis spans 16 distinct podcast domains, alongside evaluations on established datasets for Modern Greek. The findings indicate consistent WER improvements, correlating with increases in both data volume and model size. Our study confirms that assembling large, weakly supervised corpora serves as a cost-effective strategy for advancing speech technologies in under-resourced languages.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15284
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
Paraskevopoulos, Georgios
Tsoukala, Chara
Katsamanis, Athanasios
Katsouros, Vassilis
Computation and Language
Sound
Audio and Speech Processing
The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue is exacerbated in the era of large, data-intensive models. Recent research has underscored the potential of leveraging weak supervision to augment the pool of available data. In this study, we compile an 800-hour corpus of Modern Greek from podcasts and employ Whisper large-v3 to generate silver transcriptions. This corpus is utilized to fine-tune our models, aiming to assess the efficacy of this approach in enhancing ASR performance. Our analysis spans 16 distinct podcast domains, alongside evaluations on established datasets for Modern Greek. The findings indicate consistent WER improvements, correlating with increases in both data volume and model size. Our study confirms that assembling large, weakly supervised corpora serves as a cost-effective strategy for advancing speech technologies in under-resourced languages.
title The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.15284