On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Haas, Moritz, Bordt, Sebastian, von Luxburg, Ulrike, Vankadara, Leena Chennuru |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
by: Vankadara, Leena Chennuru, et al.
Published: (2026)
by: Vankadara, Leena Chennuru, et al.
Published: (2026)
μP$^2$: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling
by: Haas, Moritz, et al.
Published: (2024)
by: Haas, Moritz, et al.
Published: (2024)
How Much Can We Forget about Data Contamination?
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
Informative Post-Hoc Explanations Only Exist for Simple Functions
by: Günther, Eric, et al.
Published: (2025)
by: Günther, Eric, et al.
Published: (2025)
Rethinking Explainable Machine Learning as Applied Statistics
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
Explaining Kernel Clustering via Decision Trees
by: Fleissner, Maximilian, et al.
Published: (2024)
by: Fleissner, Maximilian, et al.
Published: (2024)
How to safely discard features based on aggregate SHAP values
by: Bhattacharjee, Robi, et al.
Published: (2025)
by: Bhattacharjee, Robi, et al.
Published: (2025)
The Manifold Hypothesis for Gradient-Based Explanations
by: Bordt, Sebastian, et al.
Published: (2022)
by: Bordt, Sebastian, et al.
Published: (2022)
Training Neural Networks at Any Scale
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Train Once, Answer All: Many Pretraining Experiments for the Cost of One
by: Bordt, Sebastian, et al.
Published: (2025)
by: Bordt, Sebastian, et al.
Published: (2025)
Using predictive multiplicity to measure individual performance within the AI Act
by: Frohnapfel, Karolin, et al.
Published: (2026)
by: Frohnapfel, Karolin, et al.
Published: (2026)
Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
The Surprising Effectiveness of Skip-Tuning in Diffusion Sampling
by: Ma, Jiajun, et al.
Published: (2024)
by: Ma, Jiajun, et al.
Published: (2024)
On Surprising Effectiveness of Masking Updates in Adaptive Optimizers
by: Joo, Taejong, et al.
Published: (2026)
by: Joo, Taejong, et al.
Published: (2026)
Self-Compatibility: Evaluating Causal Discovery without Ground Truth
by: Faller, Philipp M., et al.
Published: (2023)
by: Faller, Philipp M., et al.
Published: (2023)
Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimension
by: Haas, Moritz, et al.
Published: (2023)
by: Haas, Moritz, et al.
Published: (2023)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
by: Akyürek, Ekin, et al.
Published: (2024)
by: Akyürek, Ekin, et al.
Published: (2024)
Auditing Local Explanations is Hard
by: Bhattacharjee, Robi, et al.
Published: (2024)
by: Bhattacharjee, Robi, et al.
Published: (2024)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
by: Chen, Yilong, et al.
Published: (2026)
by: Chen, Yilong, et al.
Published: (2026)
On the Diminishing Returns of Width for Continual Learning
by: Guha, Etash, et al.
Published: (2024)
by: Guha, Etash, et al.
Published: (2024)
Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
by: Hugessen, Adriana, et al.
Published: (2024)
by: Hugessen, Adriana, et al.
Published: (2024)
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
by: Defazio, Aaron
Published: (2026)
by: Defazio, Aaron
Published: (2026)
Surprisal Driven $k$-NN for Robust and Interpretable Nonparametric Learning
by: Banerjee, Amartya, et al.
Published: (2023)
by: Banerjee, Amartya, et al.
Published: (2023)
SNOO: Step-K Nesterov Outer Optimizer - The Surprising Effectiveness of Nesterov Momentum Applied to Pseudo-Gradients
by: Kallusky, Dominik, et al.
Published: (2025)
by: Kallusky, Dominik, et al.
Published: (2025)
Virtual Width Networks
by: Seed, et al.
Published: (2025)
by: Seed, et al.
Published: (2025)
The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement Learning
by: Schneider, Moritz, et al.
Published: (2024)
by: Schneider, Moritz, et al.
Published: (2024)
Adaptive Width Neural Networks
by: Errica, Federico, et al.
Published: (2025)
by: Errica, Federico, et al.
Published: (2025)
Taming LLMs by Scaling Learning Rates with Gradient Grouping
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
World Model Robustness via Surprise Recognition
by: Zollicoffer, Geigh, et al.
Published: (2025)
by: Zollicoffer, Geigh, et al.
Published: (2025)
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
by: Zimmer, Max, et al.
Published: (2025)
by: Zimmer, Max, et al.
Published: (2025)
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
by: Mukherjee, Sagnik, et al.
Published: (2026)
by: Mukherjee, Sagnik, et al.
Published: (2026)
Disentangling Interactions and Dependencies in Feature Attribution
by: König, Gunnar, et al.
Published: (2024)
by: König, Gunnar, et al.
Published: (2024)
No More Adam: Learning Rate Scaling at Initialization is All You Need
by: Xu, Minghao, et al.
Published: (2024)
by: Xu, Minghao, et al.
Published: (2024)
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
by: Liu, Huihan, et al.
Published: (2026)
by: Liu, Huihan, et al.
Published: (2026)
Scaling Law with Learning Rate Annealing
by: Tissue, Howe, et al.
Published: (2024)
by: Tissue, Howe, et al.
Published: (2024)
Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale
by: Kaushal, Ayush, et al.
Published: (2024)
by: Kaushal, Ayush, et al.
Published: (2024)
On the Surprising Efficacy of Distillation as an Alternative to Pre-Training Small Models
by: Farhat, Sean, et al.
Published: (2024)
by: Farhat, Sean, et al.
Published: (2024)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
by: Alizadeh, Keivan, et al.
Published: (2026)
by: Alizadeh, Keivan, et al.
Published: (2026)
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
by: Sernau, Luke
Published: (2024)
by: Sernau, Luke
Published: (2024)
Les Houches Lectures on Deep Learning at Large & Infinite Width
by: Bahri, Yasaman, et al.
Published: (2023)
by: Bahri, Yasaman, et al.
Published: (2023)
Similar Items
-
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
by: Vankadara, Leena Chennuru, et al.
Published: (2026) -
μP$^2$: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling
by: Haas, Moritz, et al.
Published: (2024) -
How Much Can We Forget about Data Contamination?
by: Bordt, Sebastian, et al.
Published: (2024) -
Informative Post-Hoc Explanations Only Exist for Simple Functions
by: Günther, Eric, et al.
Published: (2025) -
Rethinking Explainable Machine Learning as Applied Statistics
by: Bordt, Sebastian, et al.
Published: (2024)