Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oberkircher, Lena S., Alabi, Jesujoba O., Klakow, Dietrich, Trouvain, Jürgen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910125943422976
author Oberkircher, Lena S.
Alabi, Jesujoba O.
Klakow, Dietrich
Trouvain, Jürgen
author_facet Oberkircher, Lena S.
Alabi, Jesujoba O.
Klakow, Dietrich
Trouvain, Jürgen
contents Natural language processing (NLP) and speech technologies have made significant progress in recent years; however, they remain largely focused on standardized language varieties. Dialects, despite their cultural significance and widespread use, are underrepresented in linguistic resources and computational models, resulting in performance disparities. To address this gap, we introduce Saar-Voice, a six-hour speech corpus for the Saarbrücken dialect of German. The dataset was created by first collecting text through digitized books and locally sourced materials. A subset of this text was recorded by nine speakers, and we conducted analyses on both the textual and speech components to assess the dataset's characteristics and quality. We discuss methodological challenges related to orthographic and speaker variation, and explore grapheme-to-phoneme (G2P) conversion. The resulting corpus provides aligned textual and audio representations. This serves as a foundation for future research on dialect-aware text-to-speech (TTS), particularly in low-resource scenarios, including zero-shot and few-shot model adaptation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11803
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
Oberkircher, Lena S.
Alabi, Jesujoba O.
Klakow, Dietrich
Trouvain, Jürgen
Computation and Language
Natural language processing (NLP) and speech technologies have made significant progress in recent years; however, they remain largely focused on standardized language varieties. Dialects, despite their cultural significance and widespread use, are underrepresented in linguistic resources and computational models, resulting in performance disparities. To address this gap, we introduce Saar-Voice, a six-hour speech corpus for the Saarbrücken dialect of German. The dataset was created by first collecting text through digitized books and locally sourced materials. A subset of this text was recorded by nine speakers, and we conducted analyses on both the textual and speech components to assess the dataset's characteristics and quality. We discuss methodological challenges related to orthographic and speaker variation, and explore grapheme-to-phoneme (G2P) conversion. The resulting corpus provides aligned textual and audio representations. This serves as a foundation for future research on dialect-aware text-to-speech (TTS), particularly in low-resource scenarios, including zero-shot and few-shot model adaptation.
title Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
topic Computation and Language
url https://arxiv.org/abs/2604.11803