Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kang, Feiyang, Ardalani, Newsha, Kuchnik, Michael, Emad, Youssef, Elhoushi, Mostafa, Sengupta, Shubhabrata, Li, Shang-Wen, Raghavendra, Ramya, Jia, Ruoxi, Wu, Carole-Jean |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
par: Kang, Feiyang, et autres
Publié: (2025)
par: Kang, Feiyang, et autres
Publié: (2025)
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
par: Mahmoud, Anas, et autres
Publié: (2023)
par: Mahmoud, Anas, et autres
Publié: (2023)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
par: Wang, Irene, et autres
Publié: (2025)
par: Wang, Irene, et autres
Publié: (2025)
To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
par: Madhyastha, Meghana, et autres
Publié: (2026)
par: Madhyastha, Meghana, et autres
Publié: (2026)
Beyond Efficiency: Scaling AI Sustainably
par: Wu, Carole-Jean, et autres
Publié: (2024)
par: Wu, Carole-Jean, et autres
Publié: (2024)
Demystifying Manifold Constraints in LLM Pre-training
par: An, Kang, et autres
Publié: (2026)
par: An, Kang, et autres
Publié: (2026)
any4: Learned 4-bit Numeric Representation for LLMs
par: Elhoushi, Mostafa, et autres
Publié: (2025)
par: Elhoushi, Mostafa, et autres
Publié: (2025)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
par: Kang, Feiyang, et autres
Publié: (2024)
par: Kang, Feiyang, et autres
Publié: (2024)
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
par: Wang, Dong, et autres
Publié: (2025)
par: Wang, Dong, et autres
Publié: (2025)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
par: Hsia, Samuel, et autres
Publié: (2023)
par: Hsia, Samuel, et autres
Publié: (2023)
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
par: Hegazy, Amr, et autres
Publié: (2025)
par: Hegazy, Amr, et autres
Publié: (2025)
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
par: Gong, Linyuan, et autres
Publié: (2024)
par: Gong, Linyuan, et autres
Publié: (2024)
ChemFM as a Scaling Law Guided Foundation Model Pre-trained on Informative Chemicals
par: Cai, Feiyang, et autres
Publié: (2024)
par: Cai, Feiyang, et autres
Publié: (2024)
Annotating the Pangenome Reveals the Diversity in the Genetic Basis for Metabolic Enzymes
par: Ardalani, Omid
Publié: (2025)
par: Ardalani, Omid
Publié: (2025)
Scaling Backwards: Minimal Synthetic Pre-training?
par: Nakamura, Ryo, et autres
Publié: (2024)
par: Nakamura, Ryo, et autres
Publié: (2024)
Scaling Laws for Pre-training Agents and World Models
par: Pearce, Tim, et autres
Publié: (2024)
par: Pearce, Tim, et autres
Publié: (2024)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
par: Kokolis, Apostolos, et autres
Publié: (2024)
par: Kokolis, Apostolos, et autres
Publié: (2024)
Composer: A Search Framework for Hybrid Neural Architecture Design
par: Acun, Bilge, et autres
Publié: (2025)
par: Acun, Bilge, et autres
Publié: (2025)
FASTTRACK: Fast and Accurate Fact Tracing for LLMs
par: Chen, Si, et autres
Publié: (2024)
par: Chen, Si, et autres
Publié: (2024)
Characterizing Model-Native Skills
par: Kang, Feiyang, et autres
Publié: (2026)
par: Kang, Feiyang, et autres
Publié: (2026)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
par: Zhang, Xinsong, et autres
Publié: (2025)
par: Zhang, Xinsong, et autres
Publié: (2025)
Structure-Aware Fill-in-the-Middle Pretraining for Code
par: Gong, Linyuan, et autres
Publié: (2025)
par: Gong, Linyuan, et autres
Publié: (2025)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
par: Gong, Linyuan, et autres
Publié: (2024)
par: Gong, Linyuan, et autres
Publié: (2024)
Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
par: Chimoto, Everlyn Asiko, et autres
Publié: (2026)
par: Chimoto, Everlyn Asiko, et autres
Publié: (2026)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
par: Zeng, Aohan, et autres
Publié: (2024)
par: Zeng, Aohan, et autres
Publié: (2024)
Exploring the Benefit of Activation Sparsity in Pre-training
par: Zhang, Zhengyan, et autres
Publié: (2024)
par: Zhang, Zhengyan, et autres
Publié: (2024)
Promises and Pitfalls of Threshold-based Auto-labeling
par: Vishwakarma, Harit, et autres
Publié: (2022)
par: Vishwakarma, Harit, et autres
Publié: (2022)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
par: Fernandez, Jared, et autres
Publié: (2024)
par: Fernandez, Jared, et autres
Publié: (2024)
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
par: Wang, Siwei, et autres
Publié: (2025)
par: Wang, Siwei, et autres
Publié: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
par: Agarwal, Saurabh, et autres
Publié: (2024)
par: Agarwal, Saurabh, et autres
Publié: (2024)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
par: Golden, Alicia, et autres
Publié: (2025)
par: Golden, Alicia, et autres
Publié: (2025)
Industrial Synthetic Segment Pre-training
par: Mae, Shinichi, et autres
Publié: (2025)
par: Mae, Shinichi, et autres
Publié: (2025)
Pre-training with Synthetic Patterns for Audio
par: Ishikawa, Yuchi, et autres
Publié: (2024)
par: Ishikawa, Yuchi, et autres
Publié: (2024)
Text Quality-Based Pruning for Efficient Training of Language Models
par: Sharma, Vasu, et autres
Publié: (2024)
par: Sharma, Vasu, et autres
Publié: (2024)
AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
par: Kang, Feiyang, et autres
Publié: (2025)
par: Kang, Feiyang, et autres
Publié: (2025)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
par: Han, Junlin, et autres
Publié: (2025)
par: Han, Junlin, et autres
Publié: (2025)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
par: Bergsma, Shane, et autres
Publié: (2025)
par: Bergsma, Shane, et autres
Publié: (2025)
Do Pre-trained Models Benefit Equally in Continual Learning?
par: Lee, Kuan-Ying, et autres
Publié: (2022)
par: Lee, Kuan-Ying, et autres
Publié: (2022)
Trusting Your AI Agent Emotionally and Cognitively: Development and Validation of a Semantic Differential Scale for AI Trust
par: Shang, Ruoxi, et autres
Publié: (2024)
par: Shang, Ruoxi, et autres
Publié: (2024)
Brevity is the soul of wit: Pruning long files for code generation
par: Singh, Aaditya K., et autres
Publié: (2024)
par: Singh, Aaditya K., et autres
Publié: (2024)
Documents similaires
-
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
par: Kang, Feiyang, et autres
Publié: (2025) -
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
par: Mahmoud, Anas, et autres
Publié: (2023) -
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
par: Wang, Irene, et autres
Publié: (2025) -
To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
par: Madhyastha, Meghana, et autres
Publié: (2026) -
Beyond Efficiency: Scaling AI Sustainably
par: Wu, Carole-Jean, et autres
Publié: (2024)