SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
Fuente:
arXiv
Saved in:
| Main Authors: | Alrashed, Sultan, Helwe, Chadi, Orabona, Francesco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs
by: Alrashed, Sultan
Published: (2024)
by: Alrashed, Sultan
Published: (2024)
Mix, MinHash, and Match: Cross-Source Agreement for Multilingual Pretraining Datasets
by: Alrashed, Sultan, et al.
Published: (2025)
by: Alrashed, Sultan, et al.
Published: (2025)
Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models
by: Alrashed, Sultan, et al.
Published: (2024)
by: Alrashed, Sultan, et al.
Published: (2024)
ATHAR: A High-Quality and Diverse Dataset for Classical Arabic to English Translation
by: Khalil, Mohammed, et al.
Published: (2024)
by: Khalil, Mohammed, et al.
Published: (2024)
MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification
by: Helwe, Chadi, et al.
Published: (2023)
by: Helwe, Chadi, et al.
Published: (2023)
High-Quality Data Augmentation for Low-Resource NMT: Combining a Translation Memory, a GAN Generator, and Filtering
by: Liu, Hengjie, et al.
Published: (2024)
by: Liu, Hengjie, et al.
Published: (2024)
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
by: Rachamalla, Neel Prabhanjan, et al.
Published: (2025)
by: Rachamalla, Neel Prabhanjan, et al.
Published: (2025)
Quality Estimation based Feedback Training for Improving Pronoun Translation
by: Dhankhar, Harshit, et al.
Published: (2025)
by: Dhankhar, Harshit, et al.
Published: (2025)
GneissWeb: Preparing High Quality Data for LLMs at Scale
by: Gohari, Hajar Emami, et al.
Published: (2025)
by: Gohari, Hajar Emami, et al.
Published: (2025)
Arabic Tweet Act: A Weighted Ensemble Pre-Trained Transformer Model for Classifying Arabic Speech Acts on Twitter
by: Alshehri, Khadejaa, et al.
Published: (2024)
by: Alshehri, Khadejaa, et al.
Published: (2024)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
by: Kim, Yungi, et al.
Published: (2024)
by: Kim, Yungi, et al.
Published: (2024)
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
by: Fleisig, Eve, et al.
Published: (2025)
by: Fleisig, Eve, et al.
Published: (2025)
BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data
by: Zou, Bo, et al.
Published: (2026)
by: Zou, Bo, et al.
Published: (2026)
Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
by: Abbas, Chaymaa, et al.
Published: (2026)
by: Abbas, Chaymaa, et al.
Published: (2026)
Post-training Large Language Models for Diverse High-Quality Responses
by: Chen, Yilei, et al.
Published: (2025)
by: Chen, Yilei, et al.
Published: (2025)
Little Giants: Synthesizing High-Quality Embedding Data at Scale
by: Chen, Haonan, et al.
Published: (2024)
by: Chen, Haonan, et al.
Published: (2024)
Scaling Parameter-Constrained Language Models with Quality Data
by: Chang, Ernie, et al.
Published: (2024)
by: Chang, Ernie, et al.
Published: (2024)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations
by: Alkhowaiter, Mohammed, et al.
Published: (2025)
by: Alkhowaiter, Mohammed, et al.
Published: (2025)
Quality-Aware Translation Tagging in Multilingual RAG system
by: Moon, Hoyeon, et al.
Published: (2025)
by: Moon, Hoyeon, et al.
Published: (2025)
Simul-LLM: A Framework for Exploring High-Quality Simultaneous Translation with Large Language Models
by: Agostinelli, Victor, et al.
Published: (2023)
by: Agostinelli, Victor, et al.
Published: (2023)
Generating High Quality Synthetic Data for Dutch Medical Conversations
by: Kuan, Cecilia, et al.
Published: (2026)
by: Kuan, Cecilia, et al.
Published: (2026)
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
by: Magdy, Samar M., et al.
Published: (2026)
by: Magdy, Samar M., et al.
Published: (2026)
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
by: Chan, Willy, et al.
Published: (2025)
by: Chan, Willy, et al.
Published: (2025)
Arctic-SnowCoder: Demystifying High-Quality Data in Code Pretraining
by: Wei, Yuxiang, et al.
Published: (2024)
by: Wei, Yuxiang, et al.
Published: (2024)
Data Quality Matters: Suicide Intention Detection on Social Media Posts Using RoBERTa-CNN
by: Lin, Emily, et al.
Published: (2024)
by: Lin, Emily, et al.
Published: (2024)
Textual Similarity as a Key Metric in Machine Translation Quality Estimation
by: Sun, Kun, et al.
Published: (2024)
by: Sun, Kun, et al.
Published: (2024)
MTQE.en-he: Machine Translation Quality Estimation for English-Hebrew
by: Rosenbaum, Andy, et al.
Published: (2026)
by: Rosenbaum, Andy, et al.
Published: (2026)
QE-EBM: Using Quality Estimators as Energy Loss for Machine Translation
by: Yoo, Gahyun, et al.
Published: (2024)
by: Yoo, Gahyun, et al.
Published: (2024)
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
by: Greisinger, Christian, et al.
Published: (2026)
by: Greisinger, Christian, et al.
Published: (2026)
Translation of Multifaceted Data without Re-Training of Machine Translation Systems
by: Moon, Hyeonseok, et al.
Published: (2024)
by: Moon, Hyeonseok, et al.
Published: (2024)
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
by: Farinhas, António, et al.
Published: (2025)
by: Farinhas, António, et al.
Published: (2025)
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
by: Pathak, Manas, et al.
Published: (2026)
by: Pathak, Manas, et al.
Published: (2026)
The Saturation Point of Backtranslation in High Quality Low Resource English Gujarati Machine Translation
by: Arif, Arwa
Published: (2025)
by: Arif, Arwa
Published: (2025)
Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model
by: Hennara, Khalil, et al.
Published: (2025)
by: Hennara, Khalil, et al.
Published: (2025)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025)
by: Ali, Mehdi, et al.
Published: (2025)
Similar Items
-
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs
by: Alrashed, Sultan
Published: (2024) -
Mix, MinHash, and Match: Cross-Source Agreement for Multilingual Pretraining Datasets
by: Alrashed, Sultan, et al.
Published: (2025) -
Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models
by: Alrashed, Sultan, et al.
Published: (2024) -
ATHAR: A High-Quality and Diverse Dataset for Classical Arabic to English Translation
by: Khalil, Mohammed, et al.
Published: (2024) -
MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification
by: Helwe, Chadi, et al.
Published: (2023)