Fantastic Pretraining Optimizers and Where to Find Them
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Kaiyue, Hall, David, Ma, Tengyu, Liang, Percy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them
by: Bui, Anh, et al.
Published: (2025)
by: Bui, Anh, et al.
Published: (2025)
Low Rank Gradients and Where to Find Them
by: Sonthalia, Rishi, et al.
Published: (2025)
by: Sonthalia, Rishi, et al.
Published: (2025)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
by: Ravichander, Abhilasha, et al.
Published: (2025)
by: Ravichander, Abhilasha, et al.
Published: (2025)
Fantastic Multi-Task Gradient Updates and How to Find Them In a Cone
by: Hassanpour, Negar, et al.
Published: (2025)
by: Hassanpour, Negar, et al.
Published: (2025)
Fantastic Biases (What are They) and Where to Find Them
by: Barriere, Valentin
Published: (2024)
by: Barriere, Valentin
Published: (2024)
Fantastic Copyrighted Beasts and How (Not) to Generate Them
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model
by: Roth, Karsten, et al.
Published: (2023)
by: Roth, Karsten, et al.
Published: (2023)
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
by: Jaiswal, Ajay, et al.
Published: (2025)
by: Jaiswal, Ajay, et al.
Published: (2025)
Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis
by: Datta, Shrestha, et al.
Published: (2026)
by: Datta, Shrestha, et al.
Published: (2026)
What Cohort INRs Encode and Where to Freeze Them
by: Sideri-Lampretsa, Vasiliki, et al.
Published: (2026)
by: Sideri-Lampretsa, Vasiliki, et al.
Published: (2026)
Calibrated Language Models and How to Find Them with Label Smoothing
by: Huang, Jerry, et al.
Published: (2025)
by: Huang, Jerry, et al.
Published: (2025)
How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
by: Su, Buxin, et al.
Published: (2025)
by: Su, Buxin, et al.
Published: (2025)
Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
by: Ramtoula, Benjamin, et al.
Published: (2025)
by: Ramtoula, Benjamin, et al.
Published: (2025)
Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)
by: Prinster, Drew, et al.
Published: (2024)
by: Prinster, Drew, et al.
Published: (2024)
Configuration-to-Performance Scaling Law with Neural Ansatz
by: Zhang, Huaqing, et al.
Published: (2026)
by: Zhang, Huaqing, et al.
Published: (2026)
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
by: Mahankali, Arvind, et al.
Published: (2026)
by: Mahankali, Arvind, et al.
Published: (2026)
STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving
by: Dong, Kefan, et al.
Published: (2025)
by: Dong, Kefan, et al.
Published: (2025)
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
by: Liu, Hong, et al.
Published: (2023)
by: Liu, Hong, et al.
Published: (2023)
Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space
by: Ruscio, Valeria, et al.
Published: (2026)
by: Ruscio, Valeria, et al.
Published: (2026)
Evaluating Self-Supervised Learning via Risk Decomposition
by: Dubois, Yann, et al.
Published: (2023)
by: Dubois, Yann, et al.
Published: (2023)
Weight Ensembling Improves Reasoning in Language Models
by: Dang, Xingyu, et al.
Published: (2025)
by: Dang, Xingyu, et al.
Published: (2025)
Reinforcement Learning for Machine Learning Engineering Agents
by: Yang, Sherry, et al.
Published: (2025)
by: Yang, Sherry, et al.
Published: (2025)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
by: Armandpour, Mohammadreza, et al.
Published: (2026)
by: Armandpour, Mohammadreza, et al.
Published: (2026)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
by: Huang, Chenghua, et al.
Published: (2025)
by: Huang, Chenghua, et al.
Published: (2025)
Linguistic Calibration of Long-Form Generations
by: Band, Neil, et al.
Published: (2024)
by: Band, Neil, et al.
Published: (2024)
Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
by: Cheng, Jie, et al.
Published: (2024)
by: Cheng, Jie, et al.
Published: (2024)
MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
by: Huang, Qian, et al.
Published: (2023)
by: Huang, Qian, et al.
Published: (2023)
Zeroth-Order Optimization Finds Flat Minima
by: Zhang, Liang, et al.
Published: (2025)
by: Zhang, Liang, et al.
Published: (2025)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
by: Zhang, Zishi, et al.
Published: (2026)
by: Zhang, Zishi, et al.
Published: (2026)
Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions
by: Ma, Nicole, et al.
Published: (2026)
by: Ma, Nicole, et al.
Published: (2026)
On the Entropy Calibration of Language Models
by: Cao, Steven, et al.
Published: (2025)
by: Cao, Steven, et al.
Published: (2025)
GNN Explanations that do not Explain and How to find Them
by: Azzolin, Steve, et al.
Published: (2026)
by: Azzolin, Steve, et al.
Published: (2026)
Large Language Models as Tool Makers
by: Cai, Tianle, et al.
Published: (2023)
by: Cai, Tianle, et al.
Published: (2023)
Spend Search Where It Pays: Value-Guided Structured Sampling and Optimization for Generative Recommendation
by: Jiang, Jie, et al.
Published: (2026)
by: Jiang, Jie, et al.
Published: (2026)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
by: Kim, Moo Jin, et al.
Published: (2025)
by: Kim, Moo Jin, et al.
Published: (2025)
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
How to Square Tensor Networks and Circuits Without Squaring Them
by: Loconte, Lorenzo, et al.
Published: (2025)
by: Loconte, Lorenzo, et al.
Published: (2025)
Similar Items
-
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025) -
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
by: Zhang, Zhenyu, et al.
Published: (2025) -
Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them
by: Bui, Anh, et al.
Published: (2025) -
Low Rank Gradients and Where to Find Them
by: Sonthalia, Rishi, et al.
Published: (2025) -
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
by: Ravichander, Abhilasha, et al.
Published: (2025)