ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
Fuente:
arXiv
Saved in:
| Main Authors: | Stankov, Vladislav, Kopp, Matyáš, Bojar, Ondřej |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
by: Barančíková, Petra, et al.
Published: (2025)
by: Barančíková, Petra, et al.
Published: (2025)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
by: Luu, Nam, et al.
Published: (2025)
by: Luu, Nam, et al.
Published: (2025)
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
by: Ljubešić, Nikola, et al.
Published: (2025)
by: Ljubešić, Nikola, et al.
Published: (2025)
Continuous Rating as Reliable Human Evaluation of Simultaneous Speech Translation
by: Javorský, Dávid, et al.
Published: (2022)
by: Javorský, Dávid, et al.
Published: (2022)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
by: Polák, Peter, et al.
Published: (2023)
by: Polák, Peter, et al.
Published: (2023)
Transfer Learning of Transformer-based Speech Recognition Models from Czech to Slovak
by: Lehečka, Jan, et al.
Published: (2023)
by: Lehečka, Jan, et al.
Published: (2023)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
by: Polák, Peter, et al.
Published: (2025)
by: Polák, Peter, et al.
Published: (2025)
Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings
by: Čechovič, Marko, et al.
Published: (2025)
by: Čechovič, Marko, et al.
Published: (2025)
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
by: Papi, Sara, et al.
Published: (2024)
by: Papi, Sara, et al.
Published: (2024)
Czech Dataset for Complex Aspect-Based Sentiment Analysis Tasks
by: Šmíd, Jakub, et al.
Published: (2025)
by: Šmíd, Jakub, et al.
Published: (2025)
Finetuning LLMs for EvaCun 2025 token prediction shared task
by: Jon, Josef, et al.
Published: (2025)
by: Jon, Josef, et al.
Published: (2025)
Understanding the role of FFNs in driving multilingual behaviour in LLMs
by: Bhattacharya, Sunit, et al.
Published: (2024)
by: Bhattacharya, Sunit, et al.
Published: (2024)
Quality and Quantity of Machine Translation References for Automatic Metrics
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet
by: Bagci, Mevlüt, et al.
Published: (2025)
by: Bagci, Mevlüt, et al.
Published: (2025)
FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
by: Teixeira, Francisco, et al.
Published: (2026)
by: Teixeira, Francisco, et al.
Published: (2026)
MEEDAV: A Synchronous Web Viewer for EEG, Eye-Tracking and Speech Data
by: Pijálek, Jan, et al.
Published: (2026)
by: Pijálek, Jan, et al.
Published: (2026)
GPT Czech Poet: Generation of Czech Poetic Strophes with Language Models
by: Chudoba, Michal, et al.
Published: (2024)
by: Chudoba, Michal, et al.
Published: (2024)
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
by: Šindelář, Pavel, et al.
Published: (2025)
by: Šindelář, Pavel, et al.
Published: (2025)
The author is dead, but what if they never lived? A reception experiment on Czech AI- and human-authored poetry
by: Marklová, Anna, et al.
Published: (2025)
by: Marklová, Anna, et al.
Published: (2025)
MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines
by: Javorský, Dávid, et al.
Published: (2025)
by: Javorský, Dávid, et al.
Published: (2025)
BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism
by: Fajcik, Martin, et al.
Published: (2024)
by: Fajcik, Martin, et al.
Published: (2024)
Prompting LLMs: Length Control for Isometric Machine Translation
by: Javorský, Dávid, et al.
Published: (2025)
by: Javorský, Dávid, et al.
Published: (2025)
Multimodal Shannon Game with Images
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
EuroSpeech: A Multilingual Speech Corpus
by: Pfisterer, Samuel, et al.
Published: (2025)
by: Pfisterer, Samuel, et al.
Published: (2025)
SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Models
by: Božík, Erik, et al.
Published: (2025)
by: Božík, Erik, et al.
Published: (2025)
CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents
by: Kostelník, Martin, et al.
Published: (2026)
by: Kostelník, Martin, et al.
Published: (2026)
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
by: Sperber, Matthias, et al.
Published: (2024)
by: Sperber, Matthias, et al.
Published: (2024)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
by: Goldin, Gili, et al.
Published: (2024)
by: Goldin, Gili, et al.
Published: (2024)
Prompt-Based Approach for Czech Sentiment Analysis
by: Šmíd, Jakub, et al.
Published: (2025)
by: Šmíd, Jakub, et al.
Published: (2025)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
by: Lehečka, Jan, et al.
Published: (2024)
by: Lehečka, Jan, et al.
Published: (2024)
RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects
by: Hassan, Md. Rezuwan, et al.
Published: (2025)
by: Hassan, Md. Rezuwan, et al.
Published: (2025)
Evaluating Optimal Reference Translations
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026)
by: Asonitis, Antonis, et al.
Published: (2026)
Hopes and Fears -- Emotion Distribution in the Topic Landscape of Finnish Parliamentary Speech 2000-2020
by: Ristilä, Anna, et al.
Published: (2026)
by: Ristilä, Anna, et al.
Published: (2026)
ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus
by: Ogunremi, Tolulope, et al.
Published: (2023)
by: Ogunremi, Tolulope, et al.
Published: (2023)
FFSTC: Fongbe to French Speech Translation Corpus
by: Kponou, D. Fortune, et al.
Published: (2024)
by: Kponou, D. Fortune, et al.
Published: (2024)
Analyzing German Parliamentary Speeches: A Machine Learning Approach for Topic and Sentiment Classification
by: Pätz, Lukas, et al.
Published: (2025)
by: Pätz, Lukas, et al.
Published: (2025)
Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus
by: Ouzerrout, Samy
Published: (2025)
by: Ouzerrout, Samy
Published: (2025)
KazParC: Kazakh Parallel Corpus for Machine Translation
by: Yeshpanov, Rustem, et al.
Published: (2024)
by: Yeshpanov, Rustem, et al.
Published: (2024)
Similar Items
-
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
by: Barančíková, Petra, et al.
Published: (2025) -
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
by: Luu, Nam, et al.
Published: (2025) -
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
by: Ljubešić, Nikola, et al.
Published: (2025) -
Continuous Rating as Reliable Human Evaluation of Simultaneous Speech Translation
by: Javorský, Dávid, et al.
Published: (2022) -
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
by: Polák, Peter, et al.
Published: (2023)