Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Zichun, Xiong, Chenyan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
by: Li, Xiaochuan, et al.
Published: (2024)
by: Li, Xiaochuan, et al.
Published: (2024)
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
by: Yu, Zichun, et al.
Published: (2024)
by: Yu, Zichun, et al.
Published: (2024)
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
Group-Level Data Selection for Efficient Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
by: Wang, Tevin, et al.
Published: (2025)
by: Wang, Tevin, et al.
Published: (2025)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
by: Kang, Hao, et al.
Published: (2025)
by: Kang, Hao, et al.
Published: (2025)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance
by: Sun, Liwen, et al.
Published: (2024)
by: Sun, Liwen, et al.
Published: (2024)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
Text2Data: Low-Resource Data Generation with Textual Control
by: Wang, Shiyu, et al.
Published: (2024)
by: Wang, Shiyu, et al.
Published: (2024)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
by: Niklaus, Joel, et al.
Published: (2026)
by: Niklaus, Joel, et al.
Published: (2026)
Many-to-English Machine Translation Tools, Data, and Pretrained Models
by: Gowda, Thamme, et al.
Published: (2021)
by: Gowda, Thamme, et al.
Published: (2021)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
by: Singh, Karan, et al.
Published: (2026)
by: Singh, Karan, et al.
Published: (2026)
MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification
by: Zheng, Bo, et al.
Published: (2026)
by: Zheng, Bo, et al.
Published: (2026)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
by: Ovalle, Anaelia, et al.
Published: (2023)
by: Ovalle, Anaelia, et al.
Published: (2023)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
by: Roger, Alexis, et al.
Published: (2025)
by: Roger, Alexis, et al.
Published: (2025)
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
by: Nguyen, Huu, et al.
Published: (2025)
by: Nguyen, Huu, et al.
Published: (2025)
Fine-tuning can Help Detect Pretraining Data from Large Language Models
by: Zhang, Hengxiang, et al.
Published: (2024)
by: Zhang, Hengxiang, et al.
Published: (2024)
PithTrain: A Compact and Agent-Native MoE Training System
by: Lai, Ruihang, et al.
Published: (2026)
by: Lai, Ruihang, et al.
Published: (2026)
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
by: Tong, Ziyi, et al.
Published: (2026)
by: Tong, Ziyi, et al.
Published: (2026)
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data
by: Kwak, Minseo, et al.
Published: (2026)
by: Kwak, Minseo, et al.
Published: (2026)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
by: Ge, Ce, et al.
Published: (2024)
by: Ge, Ce, et al.
Published: (2024)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
by: Feng, Steven, et al.
Published: (2024)
by: Feng, Steven, et al.
Published: (2024)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
by: Huang, Ruizhe, et al.
Published: (2025)
by: Huang, Ruizhe, et al.
Published: (2025)
UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science
by: Yang, Yazheng, et al.
Published: (2023)
by: Yang, Yazheng, et al.
Published: (2023)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025)
by: Ali, Mehdi, et al.
Published: (2025)
Scaling Data-Constrained Language Models
by: Muennighoff, Niklas, et al.
Published: (2023)
by: Muennighoff, Niklas, et al.
Published: (2023)
MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
by: Chen, Zhixun, et al.
Published: (2025)
by: Chen, Zhixun, et al.
Published: (2025)
Scaling Optimal LR Across Token Horizons
by: Bjorck, Johan, et al.
Published: (2024)
by: Bjorck, Johan, et al.
Published: (2024)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
by: Xiong, Weimin, et al.
Published: (2026)
by: Xiong, Weimin, et al.
Published: (2026)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions
by: Choe, Sang Keun, et al.
Published: (2024)
by: Choe, Sang Keun, et al.
Published: (2024)
Factual Inconsistency in Data-to-Text Generation Scales Exponentially with LLM Size: A Statistical Validation
by: Mahapatra, Joy, et al.
Published: (2025)
by: Mahapatra, Joy, et al.
Published: (2025)
Pretrained Generative Language Models as General Learning Frameworks for Sequence-Based Tasks
by: Fauber, Ben
Published: (2024)
by: Fauber, Ben
Published: (2024)
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
by: Niu, Tianyi, et al.
Published: (2026)
by: Niu, Tianyi, et al.
Published: (2026)
Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum
by: Song, Zihui, et al.
Published: (2026)
by: Song, Zihui, et al.
Published: (2026)
Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization
by: Yu, Jing, et al.
Published: (2025)
by: Yu, Jing, et al.
Published: (2025)
Similar Items
-
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
by: Li, Xiaochuan, et al.
Published: (2024) -
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
by: Yu, Zichun, et al.
Published: (2024) -
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
by: Yu, Zichun, et al.
Published: (2025) -
Group-Level Data Selection for Efficient Pretraining
by: Yu, Zichun, et al.
Published: (2025) -
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
by: Wang, Tevin, et al.
Published: (2025)