Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bergsma, Shane, Dey, Nolan, Hestness, Joel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling with Collapse: Efficient and Predictable Training of LLM Families
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Sparse maximal update parameterization: A holistic approach to sparse training dynamics
von: Dey, Nolan, et al.
Veröffentlicht: (2024)
von: Dey, Nolan, et al.
Veröffentlicht: (2024)
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
von: Goffinet, Etienne, et al.
Veröffentlicht: (2025)
von: Goffinet, Etienne, et al.
Veröffentlicht: (2025)
Don't be lazy: CompleteP enables compute-efficient deep transformers
von: Dey, Nolan, et al.
Veröffentlicht: (2025)
von: Dey, Nolan, et al.
Veröffentlicht: (2025)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
von: Gray, Gavia, et al.
Veröffentlicht: (2024)
von: Gray, Gavia, et al.
Veröffentlicht: (2024)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
von: Bai, Yuyang, et al.
Veröffentlicht: (2026)
von: Bai, Yuyang, et al.
Veröffentlicht: (2026)
How to Train Data-Efficient LLMs
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
ReIFE: Re-evaluating Instruction-Following Evaluation
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
von: Hase, Peter, et al.
Veröffentlicht: (2024)
von: Hase, Peter, et al.
Veröffentlicht: (2024)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
Is Child-Directed Speech Effective Training Data for Language Models?
von: Feng, Steven Y., et al.
Veröffentlicht: (2024)
von: Feng, Steven Y., et al.
Veröffentlicht: (2024)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
von: Kadhe, Swanand Ravindra, et al.
Veröffentlicht: (2024)
von: Kadhe, Swanand Ravindra, et al.
Veröffentlicht: (2024)
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
ClinicRealm: Re-evaluating Large Language Models with Conventional Machine Learning for Non-Generative Clinical Prediction Tasks
von: Zhu, Yinghao, et al.
Veröffentlicht: (2024)
von: Zhu, Yinghao, et al.
Veröffentlicht: (2024)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
von: Kang, Feiyang, et al.
Veröffentlicht: (2024)
von: Kang, Feiyang, et al.
Veröffentlicht: (2024)
Efficient multi-prompt evaluation of LLMs
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning
von: Zhang, Jianguo, et al.
Veröffentlicht: (2024)
von: Zhang, Jianguo, et al.
Veröffentlicht: (2024)
Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction
von: Li, Boyi, et al.
Veröffentlicht: (2022)
von: Li, Boyi, et al.
Veröffentlicht: (2022)
Curriculum Learning with Quality-Driven Data Selection
von: Wu, Biao, et al.
Veröffentlicht: (2024)
von: Wu, Biao, et al.
Veröffentlicht: (2024)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
Cost-Effective Hallucination Detection for LLMs
von: Valentin, Simon, et al.
Veröffentlicht: (2024)
von: Valentin, Simon, et al.
Veröffentlicht: (2024)
tinyBenchmarks: evaluating LLMs with fewer examples
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
von: Treutlein, Johannes, et al.
Veröffentlicht: (2024)
von: Treutlein, Johannes, et al.
Veröffentlicht: (2024)
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
von: Bose, Avinandan, et al.
Veröffentlicht: (2025)
von: Bose, Avinandan, et al.
Veröffentlicht: (2025)
From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
von: Jiang, Xitai, et al.
Veröffentlicht: (2026)
von: Jiang, Xitai, et al.
Veröffentlicht: (2026)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
Low-Rank Quantization-Aware Training for LLMs
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
von: Luo, Kairong, et al.
Veröffentlicht: (2025)
von: Luo, Kairong, et al.
Veröffentlicht: (2025)
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
von: Zhao, Guojiang, et al.
Veröffentlicht: (2025)
von: Zhao, Guojiang, et al.
Veröffentlicht: (2025)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
von: Guo, Song, et al.
Veröffentlicht: (2024)
von: Guo, Song, et al.
Veröffentlicht: (2024)
ReALM: Reference Resolution As Language Modeling
von: Moniz, Joel Ruben Antony, et al.
Veröffentlicht: (2024)
von: Moniz, Joel Ruben Antony, et al.
Veröffentlicht: (2024)
Zeroth-Order Adaptive Neuron Alignment Based Pruning without Re-Training
von: Cunegatti, Elia, et al.
Veröffentlicht: (2024)
von: Cunegatti, Elia, et al.
Veröffentlicht: (2024)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
von: Giordani, Jeremiah
Veröffentlicht: (2025)
von: Giordani, Jeremiah
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling with Collapse: Efficient and Predictable Training of LLM Families
von: Bergsma, Shane, et al.
Veröffentlicht: (2025) -
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
von: Bergsma, Shane, et al.
Veröffentlicht: (2025) -
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
von: Bergsma, Shane, et al.
Veröffentlicht: (2025) -
Sparse maximal update parameterization: A holistic approach to sparse training dynamics
von: Dey, Nolan, et al.
Veröffentlicht: (2024) -
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
von: Goffinet, Etienne, et al.
Veröffentlicht: (2025)