AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Feiyang, Sun, Yifan, Wen, Bingbing, Chen, Si, Song, Dawn, Mahmood, Rafid, Jia, Ruoxi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FASTTRACK: Fast and Accurate Fact Tracing for LLMs
by: Chen, Si, et al.
Published: (2024)
by: Chen, Si, et al.
Published: (2024)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
by: Mao, Shizhuo, et al.
Published: (2025)
by: Mao, Shizhuo, et al.
Published: (2025)
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
by: Wen, Bingbing, et al.
Published: (2026)
by: Wen, Bingbing, et al.
Published: (2026)
Characterizing Model-Native Skills
by: Kang, Feiyang, et al.
Published: (2026)
by: Kang, Feiyang, et al.
Published: (2026)
Data Shapley in One Training Run
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
Routing, Cascades, and User Choice for LLMs
by: Mahmood, Rafid
Published: (2026)
by: Mahmood, Rafid
Published: (2026)
Retracing the Past: LLMs Emit Training Data When They Get Lost
by: Ko, Myeongseob, et al.
Published: (2025)
by: Ko, Myeongseob, et al.
Published: (2025)
BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain Abstraction
by: Li, Jiangmeng, et al.
Published: (2024)
by: Li, Jiangmeng, et al.
Published: (2024)
Exploring the Role of Knowledge Graph-Based RAG in Japanese Medical Question Answering with Small-Scale LLMs
by: Chen, Yingjian, et al.
Published: (2025)
by: Chen, Yingjian, et al.
Published: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
by: Dimlioglu, Tolga, et al.
Published: (2026)
by: Dimlioglu, Tolga, et al.
Published: (2026)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
by: Wang, Chenguang, et al.
Published: (2024)
by: Wang, Chenguang, et al.
Published: (2024)
LLMs Can Plan Only If We Tell Them
by: Sel, Bilgehan, et al.
Published: (2025)
by: Sel, Bilgehan, et al.
Published: (2025)
A Sustainable AI Economy Needs Data Deals That Work for Generators
by: Jia, Ruoxi, et al.
Published: (2026)
by: Jia, Ruoxi, et al.
Published: (2026)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
Scaling Laws for Predicting Downstream Performance in LLMs
by: Chen, Yangyi, et al.
Published: (2024)
by: Chen, Yangyi, et al.
Published: (2024)
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
by: Li, Shengrui, et al.
Published: (2026)
by: Li, Shengrui, et al.
Published: (2026)
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
by: Song, Haiyue, et al.
Published: (2026)
by: Song, Haiyue, et al.
Published: (2026)
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
by: Chen, Jiefeng, et al.
Published: (2025)
by: Chen, Jiefeng, et al.
Published: (2025)
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
by: Goffinet, Etienne, et al.
Published: (2025)
by: Goffinet, Etienne, et al.
Published: (2025)
AutoScale: Linear Scalarization Guided by Multi-Task Optimization Metrics
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
by: Baroian, Andrei, et al.
Published: (2025)
by: Baroian, Andrei, et al.
Published: (2025)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
by: Zeng, Liang, et al.
Published: (2025)
by: Zeng, Liang, et al.
Published: (2025)
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
by: Jahan, Rafid Ishrak, et al.
Published: (2026)
by: Jahan, Rafid Ishrak, et al.
Published: (2026)
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
by: Jiang, Changhao, et al.
Published: (2025)
by: Jiang, Changhao, et al.
Published: (2025)
AutoToM: Scaling Model-based Mental Inference via Automated Agent Modeling
by: Zhang, Zhining, et al.
Published: (2025)
by: Zhang, Zhining, et al.
Published: (2025)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
by: Ceron, Tanise, et al.
Published: (2025)
by: Ceron, Tanise, et al.
Published: (2025)
From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
by: Shen, Yingli, et al.
Published: (2025)
by: Shen, Yingli, et al.
Published: (2025)
GneissWeb: Preparing High Quality Data for LLMs at Scale
by: Gohari, Hajar Emami, et al.
Published: (2025)
by: Gohari, Hajar Emami, et al.
Published: (2025)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
by: Cheng, Ruoxi, et al.
Published: (2025)
by: Cheng, Ruoxi, et al.
Published: (2025)
Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
by: Si, Jianfeng, et al.
Published: (2025)
by: Si, Jianfeng, et al.
Published: (2025)
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
AutoMix: Automatically Mixing Language Models
by: Aggarwal, Pranjal, et al.
Published: (2023)
by: Aggarwal, Pranjal, et al.
Published: (2023)
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
by: Chen, Maximillian, et al.
Published: (2024)
by: Chen, Maximillian, et al.
Published: (2024)
Scaling LLM Pre-training with Vocabulary Curriculum
by: Yu, Fangyuan
Published: (2025)
by: Yu, Fangyuan
Published: (2025)
Predicting Task Performance with Context-aware Scaling Laws
by: Montgomery, Kyle, et al.
Published: (2025)
by: Montgomery, Kyle, et al.
Published: (2025)
Similar Items
-
FASTTRACK: Fast and Accurate Fact Tracing for LLMs
by: Chen, Si, et al.
Published: (2024) -
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
by: Kang, Feiyang, et al.
Published: (2025) -
Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
by: Kang, Feiyang, et al.
Published: (2024) -
AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
by: Kang, Feiyang, et al.
Published: (2025) -
SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
by: Mao, Shizhuo, et al.
Published: (2025)