Benchmarking Optimizers for Large Language Model Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Semenov, Andrei, Pagliardini, Matteo, Jaggi, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023)
by: Fan, Simin, et al.
Published: (2023)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
Leveraging the true depth of LLMs
by: González, Ramón Calvo, et al.
Published: (2025)
by: González, Ramón Calvo, et al.
Published: (2025)
The AdEMAMix Optimizer: Better, Faster, Older
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
by: Fan, Simin, et al.
Published: (2025)
by: Fan, Simin, et al.
Published: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Stochastic Difference-of-Convex Optimization with Momentum
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
A Primal-Dual Approach to Solving Variational Inequalities with General Constraints
by: Chavdarova, Tatjana, et al.
Published: (2022)
by: Chavdarova, Tatjana, et al.
Published: (2022)
A Split-Client Approach to Second-Order Optimization
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Protein Fold Classification at Scale: Benchmarking and Pretraining
by: Chen, Dexiong, et al.
Published: (2026)
by: Chen, Dexiong, et al.
Published: (2026)
CoBo: Collaborative Learning via Bilevel Optimization
by: Hashemi, Diba, et al.
Published: (2024)
by: Hashemi, Diba, et al.
Published: (2024)
CoPeP: Benchmarking Continual Pretraining for Protein Language Models
by: Patil, Darshan, et al.
Published: (2026)
by: Patil, Darshan, et al.
Published: (2026)
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
by: Galinkin, Erick, et al.
Published: (2024)
by: Galinkin, Erick, et al.
Published: (2024)
Benchmarking Large Language Model Uncertainty for Prompt Optimization
by: Guo, Pei-Fu, et al.
Published: (2024)
by: Guo, Pei-Fu, et al.
Published: (2024)
HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
by: Kosson, Atli, et al.
Published: (2024)
by: Kosson, Atli, et al.
Published: (2024)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
by: Kosson, Atli, et al.
Published: (2023)
by: Kosson, Atli, et al.
Published: (2023)
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
AdaGC: Improving Training Stability for Large Language Model Pretraining
by: Wang, Guoxia, et al.
Published: (2025)
by: Wang, Guoxia, et al.
Published: (2025)
A New First-Order Meta-Learning Algorithm with Convergence Guarantees
by: Chayti, El Mahdi, et al.
Published: (2024)
by: Chayti, El Mahdi, et al.
Published: (2024)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
TPLLM: A Traffic Prediction Framework Based on Pretrained Large Language Models
by: Ren, Yilong, et al.
Published: (2024)
by: Ren, Yilong, et al.
Published: (2024)
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
by: Ruis, Laura, et al.
Published: (2024)
by: Ruis, Laura, et al.
Published: (2024)
Narrowing the Focus: Learned Optimizers for Pretrained Models
by: Kristiansen, Gus, et al.
Published: (2024)
by: Kristiansen, Gus, et al.
Published: (2024)
Stochastic Optimization with Random Search
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
Detecting Pretraining Data from Large Language Models
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
Data Mixing for Large Language Models Pretraining: A Survey and Outlook
by: Chen, Zhuo, et al.
Published: (2026)
by: Chen, Zhuo, et al.
Published: (2026)
A Step Toward Federated Pretraining of Multimodal Large Language Models
by: Xiong, Baochen, et al.
Published: (2026)
by: Xiong, Baochen, et al.
Published: (2026)
Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness
by: Kornilov, Nikita, et al.
Published: (2025)
by: Kornilov, Nikita, et al.
Published: (2025)
Pretrained Joint Predictions for Scalable Batch Bayesian Optimization of Molecular Designs
by: Wang-Henderson, Miles, et al.
Published: (2025)
by: Wang-Henderson, Miles, et al.
Published: (2025)
Semantic Structure of Feature Space in Large Language Models
by: Kozlowski, Austin C., et al.
Published: (2026)
by: Kozlowski, Austin C., et al.
Published: (2026)
Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
by: Portes, Jacob, et al.
Published: (2025)
by: Portes, Jacob, et al.
Published: (2025)
Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining
by: Sow, Daouda, et al.
Published: (2025)
by: Sow, Daouda, et al.
Published: (2025)
Improving Stochastic Cubic Newton with Momentum
by: Chayti, El Mahdi, et al.
Published: (2024)
by: Chayti, El Mahdi, et al.
Published: (2024)
Similar Items
-
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023) -
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023) -
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
by: Semenov, Andrei, et al.
Published: (2025) -
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024) -
Leveraging the true depth of LLMs
by: González, Ramón Calvo, et al.
Published: (2025)