ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
Fuente:
arXiv
Saved in:
| Main Authors: | Ljubešić, Nikola, Rupnik, Peter, Porupski, Ivan, Pungeršek, Taja Kuzman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
by: Pungeršek, Taja Kuzman, et al.
Published: (2025)
by: Pungeršek, Taja Kuzman, et al.
Published: (2025)
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings
by: Mochtak, Michal, et al.
Published: (2023)
by: Mochtak, Michal, et al.
Published: (2023)
Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models
by: Ljubešić, Nikola, et al.
Published: (2025)
by: Ljubešić, Nikola, et al.
Published: (2025)
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
by: Kuzman, Taja, et al.
Published: (2024)
by: Kuzman, Taja, et al.
Published: (2024)
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
by: Vintar, Špela, et al.
Published: (2025)
by: Vintar, Špela, et al.
Published: (2025)
Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024)
by: van Noord, Rik, et al.
Published: (2024)
Mići Princ -- A Little Boy Teaching Speech Technologies the Chakavian Dialect
by: Ljubešić, Nikola, et al.
Published: (2026)
by: Ljubešić, Nikola, et al.
Published: (2026)
CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
New Textual Corpora for Serbian Language Modeling
by: Škorić, Mihailo, et al.
Published: (2024)
by: Škorić, Mihailo, et al.
Published: (2024)
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
by: Stankov, Vladislav, et al.
Published: (2025)
by: Stankov, Vladislav, et al.
Published: (2025)
Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian
by: Medvecki, Darija, et al.
Published: (2024)
by: Medvecki, Darija, et al.
Published: (2024)
AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
An Annotation Scheme for Factuality and its Application to Parliamentary Proceedings
by: Goldin, Gili, et al.
Published: (2025)
by: Goldin, Gili, et al.
Published: (2025)
A Survey on Spoken Italian Datasets and Corpora
by: Giordano, Marco, et al.
Published: (2025)
by: Giordano, Marco, et al.
Published: (2025)
FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
by: Teixeira, Francisco, et al.
Published: (2026)
by: Teixeira, Francisco, et al.
Published: (2026)
Affective Polarization across European Parliaments
by: Evkoski, Bojan, et al.
Published: (2025)
by: Evkoski, Bojan, et al.
Published: (2025)
Geographic Adaptation of Pretrained Language Models
by: Hofmann, Valentin, et al.
Published: (2022)
by: Hofmann, Valentin, et al.
Published: (2022)
Test data for the shared task Ideology and Power Identification in Parliamentary Debates 2024
by: Çöltekin, Çağrı, et al.
Published: (2024)
by: Çöltekin, Çağrı, et al.
Published: (2024)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
by: Goldin, Gili, et al.
Published: (2024)
by: Goldin, Gili, et al.
Published: (2024)
Multilingual Power and Ideology Identification in the Parliament: a Reference Dataset and Simple Baselines
by: Çöltekin, Çağrı, et al.
Published: (2024)
by: Çöltekin, Çağrı, et al.
Published: (2024)
Balkan Holocausts?: Serbian and Croatian victim centred propaganda and the war in Yugoslavia
by: Macdonald, David Bruce
Published: (2010)
by: Macdonald, David Bruce
Published: (2010)
WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
by: Dai, Yuhang, et al.
Published: (2025)
by: Dai, Yuhang, et al.
Published: (2025)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
CAMEO: Collection of Multilingual Emotional Speech Corpora
by: Christop, Iwona, et al.
Published: (2025)
by: Christop, Iwona, et al.
Published: (2025)
Hopes and Fears -- Emotion Distribution in the Topic Landscape of Finnish Parliamentary Speech 2000-2020
by: Ristilä, Anna, et al.
Published: (2026)
by: Ristilä, Anna, et al.
Published: (2026)
Unsupervised Cross-Lingual Part-of-Speech Tagging with Monolingual Corpora Only
by: Zheng, Jianyu
Published: (2026)
by: Zheng, Jianyu
Published: (2026)
Building Corpora for Single-Channel Speech Separation Across Multiple Domains
by: Maciejewski, Matthew, et al.
Published: (2018)
by: Maciejewski, Matthew, et al.
Published: (2018)
Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet
by: Bagci, Mevlüt, et al.
Published: (2025)
by: Bagci, Mevlüt, et al.
Published: (2025)
Analyzing German Parliamentary Speeches: A Machine Learning Approach for Topic and Sentiment Classification
by: Pätz, Lukas, et al.
Published: (2025)
by: Pätz, Lukas, et al.
Published: (2025)
SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
by: Si, Shuzheng, et al.
Published: (2023)
by: Si, Shuzheng, et al.
Published: (2023)
A Method for Learning Large-Scale Computational Construction Grammars from Semantically Annotated Corpora
by: Van Eecke, Paul, et al.
Published: (2026)
by: Van Eecke, Paul, et al.
Published: (2026)
What Makes You CLIC: Detection of Croatian Clickbait Headlines
by: Anđelić, Marija, et al.
Published: (2025)
by: Anđelić, Marija, et al.
Published: (2025)
Scaling Spoken Language Models with Syllabic Speech Tokenization
by: Lee, Nicholas, et al.
Published: (2025)
by: Lee, Nicholas, et al.
Published: (2025)
nEMO: Dataset of Emotional Speech in Polish
by: Christop, Iwona
Published: (2024)
by: Christop, Iwona
Published: (2024)
Indigenous Languages Spoken in Argentina: A Survey of NLP and Speech Resources
by: Ticona, Belu, et al.
Published: (2025)
by: Ticona, Belu, et al.
Published: (2025)
Similar Items
-
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
by: Ljubešić, Nikola, et al.
Published: (2024) -
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
by: Pungeršek, Taja Kuzman, et al.
Published: (2026) -
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
by: Pungeršek, Taja Kuzman, et al.
Published: (2026) -
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
by: Pungeršek, Taja Kuzman, et al.
Published: (2025) -
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
by: Ljubešić, Nikola, et al.
Published: (2024)