A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Blaschke, Verena, Winkler, Miriam, Förster, Constantin, Wenger-Glemser, Gabriele, Plank, Barbara
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912610874556416
author Blaschke, Verena
Winkler, Miriam
Förster, Constantin
Wenger-Glemser, Gabriele
Plank, Barbara
author_facet Blaschke, Verena
Winkler, Miriam
Förster, Constantin
Wenger-Glemser, Gabriele
Plank, Barbara
contents Although Germany has a diverse landscape of dialects, they are underrepresented in current automatic speech recognition (ASR) research. To enable studies of how robust models are towards dialectal variation, we present Betthupferl, an evaluation dataset containing four hours of read speech in three dialect groups spoken in Southeast Germany (Franconian, Bavarian, Alemannic), and half an hour of Standard German speech. We provide both dialectal and Standard German transcriptions, and analyze the linguistic differences between them. We benchmark several multilingual state-of-the-art ASR models on speech translation into Standard German, and find differences between how much the output resembles the dialectal vs. standardized transcriptions. Qualitative error analyses of the best ASR model reveal that it sometimes normalizes grammatical differences, but often stays closer to the dialectal constructions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02894
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
Blaschke, Verena
Winkler, Miriam
Förster, Constantin
Wenger-Glemser, Gabriele
Plank, Barbara
Computation and Language
Audio and Speech Processing
Although Germany has a diverse landscape of dialects, they are underrepresented in current automatic speech recognition (ASR) research. To enable studies of how robust models are towards dialectal variation, we present Betthupferl, an evaluation dataset containing four hours of read speech in three dialect groups spoken in Southeast Germany (Franconian, Bavarian, Alemannic), and half an hour of Standard German speech. We provide both dialectal and Standard German transcriptions, and analyze the linguistic differences between them. We benchmark several multilingual state-of-the-art ASR models on speech translation into Standard German, and find differences between how much the output resembles the dialectal vs. standardized transcriptions. Qualitative error analyses of the best ASR model reveal that it sometimes normalizes grammatical differences, but often stays closer to the dialectal constructions.
title A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.02894