CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ljubešić, Nikola, Kuzman, Taja |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2026)
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2026)
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
von: Kuzman, Taja, et al.
Veröffentlicht: (2024)
von: Kuzman, Taja, et al.
Veröffentlicht: (2024)
CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2025)
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2025)
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2025)
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2025)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
von: van Noord, Rik, et al.
Veröffentlicht: (2024)
von: van Noord, Rik, et al.
Veröffentlicht: (2024)
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2026)
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2026)
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
von: Vintar, Špela, et al.
Veröffentlicht: (2025)
von: Vintar, Špela, et al.
Veröffentlicht: (2025)
Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
Wiki Dumps to Training Corpora: South Slavic Case
von: Škorić, Mihailo, et al.
Veröffentlicht: (2026)
von: Škorić, Mihailo, et al.
Veröffentlicht: (2026)
New Textual Corpora for Serbian Language Modeling
von: Škorić, Mihailo, et al.
Veröffentlicht: (2024)
von: Škorić, Mihailo, et al.
Veröffentlicht: (2024)
Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2025)
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2025)
The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings
von: Mochtak, Michal, et al.
Veröffentlicht: (2023)
von: Mochtak, Michal, et al.
Veröffentlicht: (2023)
Mići Princ -- A Little Boy Teaching Speech Technologies the Chakavian Dialect
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2026)
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2026)
Geographic Adaptation of Pretrained Language Models
von: Hofmann, Valentin, et al.
Veröffentlicht: (2022)
von: Hofmann, Valentin, et al.
Veröffentlicht: (2022)
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
Affective Polarization across European Parliaments
von: Evkoski, Bojan, et al.
Veröffentlicht: (2025)
von: Evkoski, Bojan, et al.
Veröffentlicht: (2025)
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
von: Bhattacharjee, Soham, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Soham, et al.
Veröffentlicht: (2025)
Advances in Formal Slavic Linguistics 2022
Veröffentlicht: (2025)
Veröffentlicht: (2025)
Multilingual Power and Ideology Identification in the Parliament: a Reference Dataset and Simple Baselines
von: Çöltekin, Çağrı, et al.
Veröffentlicht: (2024)
von: Çöltekin, Çağrı, et al.
Veröffentlicht: (2024)
Comparable Corpora: Opportunities for New Research Directions
von: Church, Kenneth
Veröffentlicht: (2025)
von: Church, Kenneth
Veröffentlicht: (2025)
Improving the quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristics
von: Fernando, Aloka, et al.
Veröffentlicht: (2025)
von: Fernando, Aloka, et al.
Veröffentlicht: (2025)
Building and Aligning Comparable Corpora
von: Saad, Motaz, et al.
Veröffentlicht: (2025)
von: Saad, Motaz, et al.
Veröffentlicht: (2025)
A Review of the Challenges with Massive Web-mined Corpora Used in Large Language Models Pre-Training
von: Perełkiewicz, Michał, et al.
Veröffentlicht: (2024)
von: Perełkiewicz, Michał, et al.
Veröffentlicht: (2024)
AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
von: Milička, Jiří, et al.
Veröffentlicht: (2025)
von: Milička, Jiří, et al.
Veröffentlicht: (2025)
Cross-lingual Named Entity Corpus for Slavic Languages
von: Piskorski, Jakub, et al.
Veröffentlicht: (2024)
von: Piskorski, Jakub, et al.
Veröffentlicht: (2024)
A Method for Learning Large-Scale Computational Construction Grammars from Semantically Annotated Corpora
von: Van Eecke, Paul, et al.
Veröffentlicht: (2026)
von: Van Eecke, Paul, et al.
Veröffentlicht: (2026)
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
von: Maurer, Maximilian, et al.
Veröffentlicht: (2026)
von: Maurer, Maximilian, et al.
Veröffentlicht: (2026)
Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora
von: Ranathunga, Surangika, et al.
Veröffentlicht: (2024)
von: Ranathunga, Surangika, et al.
Veröffentlicht: (2024)
The Moralization Corpus: Frame-Based Annotation and Analysis of Moralizing Speech Acts across Diverse Text Genres
von: Becker, Maria, et al.
Veröffentlicht: (2025)
von: Becker, Maria, et al.
Veröffentlicht: (2025)
Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring
von: Song, Jayoung, et al.
Veröffentlicht: (2025)
von: Song, Jayoung, et al.
Veröffentlicht: (2025)
Data Caricatures: On the Representation of African American Language in Pretraining Corpora
von: Deas, Nicholas, et al.
Veröffentlicht: (2025)
von: Deas, Nicholas, et al.
Veröffentlicht: (2025)
A Comparative Approach to Assessing Linguistic Creativity of Large Language Models and Humans
von: Dinu, Anca, et al.
Veröffentlicht: (2025)
von: Dinu, Anca, et al.
Veröffentlicht: (2025)
Automated Annotation of Evolving Corpora for Augmenting Longitudinal Network Data: A Framework Integrating Large Language Models and Expert Knowledge
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
Using LLMs to Aid Annotation and Collection of Clinically-Enriched Data in Bipolar Disorder and Schizophrenia
von: Aich, Ankit, et al.
Veröffentlicht: (2024)
von: Aich, Ankit, et al.
Veröffentlicht: (2024)
JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
von: Saadi, Hossain Shaikh, et al.
Veröffentlicht: (2025)
von: Saadi, Hossain Shaikh, et al.
Veröffentlicht: (2025)
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
From Outliers to Topics in Language Models: Anticipating Trends in News Corpora
von: Zve, Evangelia, et al.
Veröffentlicht: (2025)
von: Zve, Evangelia, et al.
Veröffentlicht: (2025)
Measuring Grammatical Diversity from Small Corpora: Derivational Entropy Rates, Mean Length of Utterances, and Annotation Invariance
von: Martin, Fermin Moscoso del Prado
Veröffentlicht: (2024)
von: Martin, Fermin Moscoso del Prado
Veröffentlicht: (2024)
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2026) -
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
von: Kuzman, Taja, et al.
Veröffentlicht: (2024) -
CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024) -
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2025) -
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2025)