Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Messmer, Bettina, Sabolčec, Vinko, Jaggi, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026)
by: Turki, Yassine, et al.
Published: (2026)
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Towards an empirical understanding of MoE design choices
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
by: Penedo, Guilherme, et al.
Published: (2025)
by: Penedo, Guilherme, et al.
Published: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
by: Kosson, Atli, et al.
Published: (2024)
by: Kosson, Atli, et al.
Published: (2024)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
by: Kosson, Atli, et al.
Published: (2023)
by: Kosson, Atli, et al.
Published: (2023)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Factual Consistency of Multilingual Pretrained Language Models
by: Fierro, Constanza, et al.
Published: (2022)
by: Fierro, Constanza, et al.
Published: (2022)
MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
by: Chen, Zhixun, et al.
Published: (2025)
by: Chen, Zhixun, et al.
Published: (2025)
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
by: Yu, Zichun, et al.
Published: (2024)
by: Yu, Zichun, et al.
Published: (2024)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
Group-Level Data Selection for Efficient Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning
by: Wang, Qibin, et al.
Published: (2024)
by: Wang, Qibin, et al.
Published: (2024)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025)
by: Ali, Mehdi, et al.
Published: (2025)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
by: Feng, Steven, et al.
Published: (2024)
by: Feng, Steven, et al.
Published: (2024)
Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages
by: Zhao, Yue, et al.
Published: (2026)
by: Zhao, Yue, et al.
Published: (2026)
LLM Pretraining with Continuous Concepts
by: Tack, Jihoon, et al.
Published: (2025)
by: Tack, Jihoon, et al.
Published: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023)
by: Fan, Simin, et al.
Published: (2023)
Strong Teacher Not Needed? On Distillation in LLM Pretraining
by: Lu, Taiming, et al.
Published: (2026)
by: Lu, Taiming, et al.
Published: (2026)
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
by: Ishibashi, Yoichi, et al.
Published: (2025)
by: Ishibashi, Yoichi, et al.
Published: (2025)
Procedural Pretraining: Warming Up Language Models with Abstract Data
by: Jiang, Liangze, et al.
Published: (2026)
by: Jiang, Liangze, et al.
Published: (2026)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
Greedy Information Projection for LLM Data Selection
by: Dong, Victor Ye, et al.
Published: (2026)
by: Dong, Victor Ye, et al.
Published: (2026)
Dissecting Outlier Dynamics in LLM NVFP4 Pretraining
by: Dong, Peijie, et al.
Published: (2026)
by: Dong, Peijie, et al.
Published: (2026)
Language Models Improve When Pretraining Data Matches Target Tasks
by: Mizrahi, David, et al.
Published: (2025)
by: Mizrahi, David, et al.
Published: (2025)
Temporal Entailment Pretraining for Clinical Language Models over EHR Data
by: Tanaka, Tatsunori, et al.
Published: (2025)
by: Tanaka, Tatsunori, et al.
Published: (2025)
Data Mixing for Large Language Models Pretraining: A Survey and Outlook
by: Chen, Zhuo, et al.
Published: (2026)
by: Chen, Zhuo, et al.
Published: (2026)
URIEL+: Enhancing Linguistic Inclusion and Usability in a Typological and Multilingual Knowledge Base
by: Khan, Aditya, et al.
Published: (2024)
by: Khan, Aditya, et al.
Published: (2024)
Leveraging the true depth of LLMs
by: González, Ramón Calvo, et al.
Published: (2025)
by: González, Ramón Calvo, et al.
Published: (2025)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
by: Ma, Xuezhe, et al.
Published: (2024)
by: Ma, Xuezhe, et al.
Published: (2024)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
by: Xie, Wanyun, et al.
Published: (2025)
by: Xie, Wanyun, et al.
Published: (2025)
Similar Items
-
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026) -
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
by: Fan, Dongyang, et al.
Published: (2025) -
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024) -
Towards an empirical understanding of MoE design choices
by: Fan, Dongyang, et al.
Published: (2024) -
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)