Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Almeida, Thales Sales, Nogueira, Rodrigo, Pedrini, Hélio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
Sabiá-2: A New Generation of Portuguese Large Language Models
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2024)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2024)
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
Measuring Cross-lingual Transfer in Bytes
von: de Souza, Leandro Rodrigues, et al.
Veröffentlicht: (2024)
von: de Souza, Leandro Rodrigues, et al.
Veröffentlicht: (2024)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
von: Bonás, Giovana Kerche, et al.
Veröffentlicht: (2026)
von: Bonás, Giovana Kerche, et al.
Veröffentlicht: (2026)
ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language
von: Piau, Marcos, et al.
Veröffentlicht: (2024)
von: Piau, Marcos, et al.
Veröffentlicht: (2024)
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
von: Junior, Roseval Malaquias, et al.
Veröffentlicht: (2026)
von: Junior, Roseval Malaquias, et al.
Veröffentlicht: (2026)
Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
von: Abonizio, Hugo, et al.
Veröffentlicht: (2025)
von: Abonizio, Hugo, et al.
Veröffentlicht: (2025)
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
SurveySum: A Dataset for Summarizing Multiple Scientific Articles into a Survey Section
von: Fernandes, Leandro Carísio, et al.
Veröffentlicht: (2024)
von: Fernandes, Leandro Carísio, et al.
Veröffentlicht: (2024)
The interplay between domain specialization and model size
von: Junior, Roseval Malaquias, et al.
Veröffentlicht: (2025)
von: Junior, Roseval Malaquias, et al.
Veröffentlicht: (2025)
Sabiá-3 Technical Report
von: Abonizio, Hugo, et al.
Veröffentlicht: (2024)
von: Abonizio, Hugo, et al.
Veröffentlicht: (2024)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2026)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2026)
Sabiá-4 Technical Report
von: Laitz, Thiago, et al.
Veröffentlicht: (2026)
von: Laitz, Thiago, et al.
Veröffentlicht: (2026)
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
von: Santos, João Guilherme Alves, et al.
Veröffentlicht: (2025)
von: Santos, João Guilherme Alves, et al.
Veröffentlicht: (2025)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
von: Nogueira, Rodrigo, et al.
Veröffentlicht: (2026)
von: Nogueira, Rodrigo, et al.
Veröffentlicht: (2026)
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
von: Nogueira, Rodrigo, et al.
Veröffentlicht: (2026)
von: Nogueira, Rodrigo, et al.
Veröffentlicht: (2026)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
von: Pires, Ramon, et al.
Veröffentlicht: (2026)
von: Pires, Ramon, et al.
Veröffentlicht: (2026)
Predictive Authoring for Brazilian Portuguese Augmentative and Alternative Communication
von: Pereira, Jayr, et al.
Veröffentlicht: (2023)
von: Pereira, Jayr, et al.
Veröffentlicht: (2023)
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2025)
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2025)
Towards High-Fidelity Synthetic Multi-platform Social Media Datasets via Large Language Models
von: Tari, Henry, et al.
Veröffentlicht: (2025)
von: Tari, Henry, et al.
Veröffentlicht: (2025)
Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers
von: Zhang, Tianhua, et al.
Veröffentlicht: (2024)
von: Zhang, Tianhua, et al.
Veröffentlicht: (2024)
Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting
von: Ho, Chloe, et al.
Veröffentlicht: (2025)
von: Ho, Chloe, et al.
Veröffentlicht: (2025)
ManufactuBERT: Efficient Continual Pretraining for Manufacturing
von: Armingaud, Robin, et al.
Veröffentlicht: (2025)
von: Armingaud, Robin, et al.
Veröffentlicht: (2025)
PICOs-RAG: PICO-supported Query Rewriting for Retrieval-Augmented Generation in Evidence-Based Medicine
von: Sun, Mengzhou, et al.
Veröffentlicht: (2025)
von: Sun, Mengzhou, et al.
Veröffentlicht: (2025)
LLM Pretraining with Continuous Concepts
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
InstaSynth: Opportunities and Challenges in Generating Synthetic Instagram Data with ChatGPT for Sponsored Content Detection
von: Bertaglia, Thales, et al.
Veröffentlicht: (2024)
von: Bertaglia, Thales, et al.
Veröffentlicht: (2024)
Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT
von: Nguyen, Duy Anh
Veröffentlicht: (2026)
von: Nguyen, Duy Anh
Veröffentlicht: (2026)
ExaRanker-Open: Synthetic Explanation for IR using Open-Source LLMs
von: Ferraretto, Fernando, et al.
Veröffentlicht: (2024)
von: Ferraretto, Fernando, et al.
Veröffentlicht: (2024)
Can Synthetic Query Rewrites Capture User Intent Better than Humans in Retrieval-Augmented Generation?
von: Zheng, JiaYing, et al.
Veröffentlicht: (2025)
von: Zheng, JiaYing, et al.
Veröffentlicht: (2025)
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
von: DatologyAI, et al.
Veröffentlicht: (2025)
von: DatologyAI, et al.
Veröffentlicht: (2025)
Yet another algorithmic bias: A Discursive Analysis of Large Language Models Reinforcing Dominant Discourses on Gender and Race
von: Bonil, Gustavo, et al.
Veröffentlicht: (2025)
von: Bonil, Gustavo, et al.
Veröffentlicht: (2025)
Enhancing Portuguese Variety Identification with Cross-Domain Approaches
von: Sousa, Hugo, et al.
Veröffentlicht: (2025)
von: Sousa, Hugo, et al.
Veröffentlicht: (2025)
LinGO: A Linguistic Graph Optimization Framework with LLMs for Interpreting Intents of Online Uncivil Discourse
von: Zhang, Yuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yuan, et al.
Veröffentlicht: (2026)
Investigating Continual Pretraining in Large Language Models: Insights and Implications
von: Yıldız, Çağatay, et al.
Veröffentlicht: (2024)
von: Yıldız, Çağatay, et al.
Veröffentlicht: (2024)
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
von: Negoita, Vlad, et al.
Veröffentlicht: (2025)
von: Negoita, Vlad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025) -
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025) -
PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025) -
Sabiá-2: A New Generation of Portuguese Large Language Models
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2024) -
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)