Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
Fuente:
arXiv
Saved in:
| Main Authors: | Almeida, Thales Sales, Nogueira, Rodrigo, Pedrini, Helio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024)
by: Almeida, Thales Sales, et al.
Published: (2024)
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026)
by: Junior, Roseval Malaquias, et al.
Published: (2026)
Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
by: Abonizio, Hugo, et al.
Published: (2025)
by: Abonizio, Hugo, et al.
Published: (2025)
CC-GPX: Extracting High-Quality Annotated Geospatial Data from Common Crawl
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
Measuring Cross-lingual Transfer in Bytes
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)
by: Bonás, Giovana Kerche, et al.
Published: (2026)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024)
by: van Noord, Rik, et al.
Published: (2024)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
by: Tessema, Bethel Melesse, et al.
Published: (2024)
by: Tessema, Bethel Melesse, et al.
Published: (2024)
SurveySum: A Dataset for Summarizing Multiple Scientific Articles into a Survey Section
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
by: Su, Dan, et al.
Published: (2024)
by: Su, Dan, et al.
Published: (2024)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
Quantifying Geospatial in the Common Crawl Corpus
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
Leveraging Web-Crawled Data for High-Quality Fine-Tuning
by: Zhou, Jing, et al.
Published: (2024)
by: Zhou, Jing, et al.
Published: (2024)
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026)
by: Pires, Ramon, et al.
Published: (2026)
The interplay between domain specialization and model size
by: Junior, Roseval Malaquias, et al.
Published: (2025)
by: Junior, Roseval Malaquias, et al.
Published: (2025)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Building and Aligning Comparable Corpora
by: Saad, Motaz, et al.
Published: (2025)
by: Saad, Motaz, et al.
Published: (2025)
Sabiá-3 Technical Report
by: Abonizio, Hugo, et al.
Published: (2024)
by: Abonizio, Hugo, et al.
Published: (2024)
Web Page Classification using LLMs for Crawling Support
by: Sasazawa, Yuichi, et al.
Published: (2025)
by: Sasazawa, Yuichi, et al.
Published: (2025)
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
by: Artemova, Ekaterina, et al.
Published: (2025)
by: Artemova, Ekaterina, et al.
Published: (2025)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
by: Hilasaca, Kenji, et al.
Published: (2026)
by: Hilasaca, Kenji, et al.
Published: (2026)
Infini-News: Efficiently Queryable Access to 1.3 Billion Processed Common Crawl News Articles
by: Lazzaroni, Ruggero Marino, et al.
Published: (2026)
by: Lazzaroni, Ruggero Marino, et al.
Published: (2026)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Sabiá-4 Technical Report
by: Laitz, Thiago, et al.
Published: (2026)
by: Laitz, Thiago, et al.
Published: (2026)
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
by: Santos, João Guilherme Alves, et al.
Published: (2025)
by: Santos, João Guilherme Alves, et al.
Published: (2025)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Building Corpora for Single-Channel Speech Separation Across Multiple Domains
by: Maciejewski, Matthew, et al.
Published: (2018)
by: Maciejewski, Matthew, et al.
Published: (2018)
From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
by: Shen, Yingli, et al.
Published: (2025)
by: Shen, Yingli, et al.
Published: (2025)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
LinGO: A Linguistic Graph Optimization Framework with LLMs for Interpreting Intents of Online Uncivil Discourse
by: Zhang, Yuan, et al.
Published: (2026)
by: Zhang, Yuan, et al.
Published: (2026)
Predictive Authoring for Brazilian Portuguese Augmentative and Alternative Communication
by: Pereira, Jayr, et al.
Published: (2023)
by: Pereira, Jayr, et al.
Published: (2023)
Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora
by: Ranathunga, Surangika, et al.
Published: (2024)
by: Ranathunga, Surangika, et al.
Published: (2024)
Similar Items
-
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2026) -
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2025) -
PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
by: Almeida, Thales Sales, et al.
Published: (2025) -
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024) -
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
by: Almeida, Thales Sales, et al.
Published: (2025)