ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Stankov, Vladislav, Kopp, Matyáš, Bojar, Ondřej
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918137517047808
author Stankov, Vladislav
Kopp, Matyáš
Bojar, Ondřej
author_facet Stankov, Vladislav
Kopp, Matyáš
Bojar, Ondřej
contents We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the sound recordings of the Czech parliamentary speeches with the official transcripts. The recordings were processed with WhisperX and Wav2Vec 2.0 to extract automated audio-text alignment. Our processing pipeline improves upon the ParCzech 3.0 speech recognition version by extracting more data with higher alignment reliability. The dataset is offered in three flexible variants: (1) sentence-segmented for automatic speech recognition and speech synthesis tasks with clean boundaries, (2) unsegmented preserving original utterance flow across sentences, and (3) a raw-alignment for further custom refinement for other possible tasks. All variants maintain the original metadata and are released under a permissive CC-BY license. The dataset is available in the LINDAT repository, with the sentence-segmented and unsegmented variants additionally available on Hugging Face.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06675
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
Stankov, Vladislav
Kopp, Matyáš
Bojar, Ondřej
Computation and Language
We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the sound recordings of the Czech parliamentary speeches with the official transcripts. The recordings were processed with WhisperX and Wav2Vec 2.0 to extract automated audio-text alignment. Our processing pipeline improves upon the ParCzech 3.0 speech recognition version by extracting more data with higher alignment reliability. The dataset is offered in three flexible variants: (1) sentence-segmented for automatic speech recognition and speech synthesis tasks with clean boundaries, (2) unsegmented preserving original utterance flow across sentences, and (3) a raw-alignment for further custom refinement for other possible tasks. All variants maintain the original metadata and are released under a permissive CC-BY license. The dataset is available in the LINDAT repository, with the sentence-segmented and unsegmented variants additionally available on Hugging Face.
title ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
topic Computation and Language
url https://arxiv.org/abs/2509.06675