SwissGPC v1.0 -- The Swiss German Podcasts Corpus

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stucki, Samuel, Cieliebak, Mark, Deriu, Jan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909804120768512
author Stucki, Samuel
Cieliebak, Mark
Deriu, Jan
author_facet Stucki, Samuel
Cieliebak, Mark
Deriu, Jan
contents We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to talk shows and podcasts hosted on Schweizer Radio und Fernsehen and YouTube, which contain approximately 5400 hours of raw audio. After segmentation and weak annotation, nearly 5000 hours of speech were retained, covering the seven major Swiss German dialect regions alongside Standard German. We describe the corpus construction methodology, including an automated annotation pipeline, and provide statistics on dialect distribution, token counts, and segmentation characteristics. Unlike existing Swiss German speech corpora, which primarily feature controlled speech, this corpus captures natural, spontaneous conversations, making it a valuable resource for real-world speech applications.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19866
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SwissGPC v1.0 -- The Swiss German Podcasts Corpus
Stucki, Samuel
Cieliebak, Mark
Deriu, Jan
Computation and Language
We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to talk shows and podcasts hosted on Schweizer Radio und Fernsehen and YouTube, which contain approximately 5400 hours of raw audio. After segmentation and weak annotation, nearly 5000 hours of speech were retained, covering the seven major Swiss German dialect regions alongside Standard German. We describe the corpus construction methodology, including an automated annotation pipeline, and provide statistics on dialect distribution, token counts, and segmentation characteristics. Unlike existing Swiss German speech corpora, which primarily feature controlled speech, this corpus captures natural, spontaneous conversations, making it a valuable resource for real-world speech applications.
title SwissGPC v1.0 -- The Swiss German Podcasts Corpus
topic Computation and Language
url https://arxiv.org/abs/2509.19866