Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | van Noord, Rik, Kuzman, Taja, Rupnik, Peter, Ljubešić, Nikola, Esplà-Gomis, Miquel, Ramírez-Sánchez, Gema, Toral, Antonio |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
par: Ljubešić, Nikola, et autres
Publié: (2024)
par: Ljubešić, Nikola, et autres
Publié: (2024)
Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining
par: Ljubešić, Nikola, et autres
Publié: (2024)
par: Ljubešić, Nikola, et autres
Publié: (2024)
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
par: Ljubešić, Nikola, et autres
Publié: (2025)
par: Ljubešić, Nikola, et autres
Publié: (2025)
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
par: Kuzman, Taja, et autres
Publié: (2024)
par: Kuzman, Taja, et autres
Publié: (2024)
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
par: Pungeršek, Taja Kuzman, et autres
Publié: (2025)
par: Pungeršek, Taja Kuzman, et autres
Publié: (2025)
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models
par: Ljubešić, Nikola, et autres
Publié: (2025)
par: Ljubešić, Nikola, et autres
Publié: (2025)
CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route
par: Ljubešić, Nikola, et autres
Publié: (2024)
par: Ljubešić, Nikola, et autres
Publié: (2024)
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
par: Vintar, Špela, et autres
Publié: (2025)
par: Vintar, Špela, et autres
Publié: (2025)
Smart Bilingual Focused Crawling of Parallel Documents
par: García-Romero, Cristian, et autres
Publié: (2024)
par: García-Romero, Cristian, et autres
Publié: (2024)
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
par: Ljubešić, Nikola, et autres
Publié: (2024)
par: Ljubešić, Nikola, et autres
Publié: (2024)
A Simple Approach to Use Bilingual Information Sources for Word Alignment
par: Miquel Esplà-Gomis
Publié: (2012)
par: Miquel Esplà-Gomis
Publié: (2012)
Mići Princ -- A Little Boy Teaching Speech Technologies the Chakavian Dialect
par: Ljubešić, Nikola, et autres
Publié: (2026)
par: Ljubešić, Nikola, et autres
Publié: (2026)
The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings
par: Mochtak, Michal, et autres
Publié: (2023)
par: Mochtak, Michal, et autres
Publié: (2023)
Multi-perspective Alignment for Increasing Naturalness in Neural Machine Translation
par: Lai, Huiyuan, et autres
Publié: (2024)
par: Lai, Huiyuan, et autres
Publié: (2024)
Towards Tailored Recovery of Lexical Diversity in Literary Machine Translation
par: Ploeger, Esther, et autres
Publié: (2024)
par: Ploeger, Esther, et autres
Publié: (2024)
Bitexor, un cosechador automático de memorias de traducción a partir de sitios web bilingües
par: Miquel Esplà
Publié: (2009)
par: Miquel Esplà
Publié: (2009)
Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
par: García-Romero, Cristian, et autres
Publié: (2025)
par: García-Romero, Cristian, et autres
Publié: (2025)
Non-Fluent Synthetic Target-Language Data Improve Neural Machine Translation
par: Sánchez-Cartagena, Víctor M., et autres
Publié: (2024)
par: Sánchez-Cartagena, Víctor M., et autres
Publié: (2024)
New Textual Corpora for Serbian Language Modeling
par: Škorić, Mihailo, et autres
Publié: (2024)
par: Škorić, Mihailo, et autres
Publié: (2024)
Document Quality Scoring for Web Crawling
par: Pezzuti, Francesca, et autres
Publié: (2025)
par: Pezzuti, Francesca, et autres
Publié: (2025)
Geographic Adaptation of Pretrained Language Models
par: Hofmann, Valentin, et autres
Publié: (2022)
par: Hofmann, Valentin, et autres
Publié: (2022)
Facts Do Care About Your Language: Assessing Answer Quality of Multilingual LLMs
par: Kansal, Yuval, et autres
Publié: (2025)
par: Kansal, Yuval, et autres
Publié: (2025)
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
par: Almeida, Thales Sales, et autres
Publié: (2025)
par: Almeida, Thales Sales, et autres
Publié: (2025)
Leveraging Web-Crawled Data for High-Quality Fine-Tuning
par: Zhou, Jing, et autres
Publié: (2024)
par: Zhou, Jing, et autres
Publié: (2024)
Cross-lingual neural fuzzy matching for exploiting target-language monolingual corpora in computer-aided translation
par: Esplà-Gomis, Miquel, et autres
Publié: (2024)
par: Esplà-Gomis, Miquel, et autres
Publié: (2024)
Neural Prioritisation for Web Crawling
par: Pezzuti, Francesca, et autres
Publié: (2025)
par: Pezzuti, Francesca, et autres
Publié: (2025)
Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation
par: Edman, Lukas, et autres
Publié: (2023)
par: Edman, Lukas, et autres
Publié: (2023)
Movie Recommendation using Web Crawling
par: Raj, Pronit, et autres
Publié: (2024)
par: Raj, Pronit, et autres
Publié: (2024)
ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation
par: Liu, Ruixuan, et autres
Publié: (2024)
par: Liu, Ruixuan, et autres
Publié: (2024)
Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations
par: Gerrits, Kyo, et autres
Publié: (2026)
par: Gerrits, Kyo, et autres
Publié: (2026)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
par: Tessema, Bethel Melesse, et autres
Publié: (2024)
par: Tessema, Bethel Melesse, et autres
Publié: (2024)
On cubic rainbow domination regular graphs
par: Kuzman, Bostjan
Publié: (2024)
par: Kuzman, Bostjan
Publié: (2024)
PMB5: Gaining More Insight into Neural Semantic Parsing with Challenging Benchmarks
par: Zhang, Xiao, et autres
Publié: (2024)
par: Zhang, Xiao, et autres
Publié: (2024)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
par: Kim, Yungi, et autres
Publié: (2024)
par: Kim, Yungi, et autres
Publié: (2024)
Bases moleculares de la carcinogénesis viral de papiloma y polioma
par: Lucía Taja
Publié: (1996)
par: Lucía Taja
Publié: (1996)
Affective Polarization across European Parliaments
par: Evkoski, Bojan, et autres
Publié: (2025)
par: Evkoski, Bojan, et autres
Publié: (2025)
Zarabanda franquista / Carlos Espla ; prólogo de Indalecio Prieto
par: Espla, Carlos
Publié: (1944)
par: Espla, Carlos
Publié: (1944)
Web Page Classification using LLMs for Crawling Support
par: Sasazawa, Yuichi, et autres
Publié: (2025)
par: Sasazawa, Yuichi, et autres
Publié: (2025)
Documents similaires
-
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026) -
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
par: Ljubešić, Nikola, et autres
Publié: (2024) -
Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining
par: Ljubešić, Nikola, et autres
Publié: (2024) -
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
par: Ljubešić, Nikola, et autres
Publié: (2025) -
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
par: Kuzman, Taja, et autres
Publié: (2024)