Per-example gradients: a new frontier for understanding and improving optimizers
Fuente:
arXiv
Saved in:
| Main Authors: | Roulet, Vincent, Agarwala, Atish |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Interplay Between Stepsize Tuning and Progressive Sharpening
by: Roulet, Vincent, et al.
Published: (2023)
by: Roulet, Vincent, et al.
Published: (2023)
High dimensional theory of two-phase optimizers
by: Agarwala, Atish
Published: (2026)
by: Agarwala, Atish
Published: (2026)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
by: Roulet, Vincent, et al.
Published: (2024)
by: Roulet, Vincent, et al.
Published: (2024)
How far away are truly hyperparameter-free learning algorithms?
by: Kasimbeg, Priya, et al.
Published: (2025)
by: Kasimbeg, Priya, et al.
Published: (2025)
What do near-optimal learning rate schedules look like?
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
by: Agarwala, Atish, et al.
Published: (2024)
by: Agarwala, Atish, et al.
Published: (2024)
Neglected Hessian component explains mysteries in Sharpness regularization
by: Dauphin, Yann N., et al.
Published: (2024)
by: Dauphin, Yann N., et al.
Published: (2024)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
by: Xiao, Ke Liang, et al.
Published: (2024)
by: Xiao, Ke Liang, et al.
Published: (2024)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
by: Qiu, Shikai, et al.
Published: (2025)
by: Qiu, Shikai, et al.
Published: (2025)
Calibration improves detection of mislabeled examples
by: Chibane, Ilies, et al.
Published: (2025)
by: Chibane, Ilies, et al.
Published: (2025)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
The Elements of Differentiable Programming
by: Blondel, Mathieu, et al.
Published: (2024)
by: Blondel, Mathieu, et al.
Published: (2024)
Joint Learning of Energy-based Models and their Partition Function
by: Sander, Michael E., et al.
Published: (2025)
by: Sander, Michael E., et al.
Published: (2025)
Near-optimal Per-Action Regret Bounds for Sleeping Bandits
by: Nguyen, Quan, et al.
Published: (2024)
by: Nguyen, Quan, et al.
Published: (2024)
Loss Functions and Operators Generated by f-Divergences
by: Roulet, Vincent, et al.
Published: (2025)
by: Roulet, Vincent, et al.
Published: (2025)
On improving generalization in a class of learning problems with the method of small parameters for weakly-controlled optimal gradient systems
by: Befekadu, Getachew K.
Published: (2024)
by: Befekadu, Getachew K.
Published: (2024)
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
by: Blondel, Mathieu, et al.
Published: (2025)
by: Blondel, Mathieu, et al.
Published: (2025)
Projected gradient methods for nonconvex and stochastic smooth optimization: new complexities and auto-conditioned stepsizes
by: Lan, Guanghui, et al.
Published: (2024)
by: Lan, Guanghui, et al.
Published: (2024)
Some remarks on gradient dominance and LQR policy optimization
by: Sontag, Eduardo D.
Published: (2025)
by: Sontag, Eduardo D.
Published: (2025)
Mislabeled examples detection viewed as probing machine learning models: concepts, survey and extensive benchmark
by: George, Thomas, et al.
Published: (2024)
by: George, Thomas, et al.
Published: (2024)
A policy gradient approach for optimization of smooth risk measures
by: Vijayan, Nithia, et al.
Published: (2022)
by: Vijayan, Nithia, et al.
Published: (2022)
A stochastic gradient method for trilevel optimization
by: Giovannelli, Tommaso, et al.
Published: (2025)
by: Giovannelli, Tommaso, et al.
Published: (2025)
A Closed-Form Persistence-Landmark Pipeline for Certified Point-Cloud and Graph Classification
by: Majhi, Sushovan, et al.
Published: (2026)
by: Majhi, Sushovan, et al.
Published: (2026)
Overshoot: Taking advantage of future gradients in momentum-based stochastic optimization
by: Kopal, Jakub, et al.
Published: (2025)
by: Kopal, Jakub, et al.
Published: (2025)
Dealing with unbounded gradients in stochastic saddle-point optimization
by: Neu, Gergely, et al.
Published: (2024)
by: Neu, Gergely, et al.
Published: (2024)
Pareto-frontier Entropy Search with Variational Lower Bound Maximization
by: Ishikura, Masanori, et al.
Published: (2025)
by: Ishikura, Masanori, et al.
Published: (2025)
Leveraging Per-Instance Privacy for Machine Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Topological Characterization of Churn Flow and Unsupervised Correction to the Wu Flow-Regime Map in Small-Diameter Vertical Pipes
by: Koenig, Brady, et al.
Published: (2026)
by: Koenig, Brady, et al.
Published: (2026)
A novel gradient-based method for decision trees optimizing arbitrary differential loss functions
by: Konstantinov, Andrei V., et al.
Published: (2025)
by: Konstantinov, Andrei V., et al.
Published: (2025)
What's the next frontier for Data-centric AI? Data Savvy Agents
by: Seedat, Nabeel, et al.
Published: (2025)
by: Seedat, Nabeel, et al.
Published: (2025)
Unregularized limit of stochastic gradient method for Wasserstein distributionally robust optimization
by: Le, Tam
Published: (2025)
by: Le, Tam
Published: (2025)
Iterative Linear Quadratic Optimization for Nonlinear Control: Differentiable Programming Algorithmic Templates
by: Roulet, Vincent, et al.
Published: (2022)
by: Roulet, Vincent, et al.
Published: (2022)
Exploring gauge-fixing conditions with gradient-based optimization
by: Detmold, William, et al.
Published: (2024)
by: Detmold, William, et al.
Published: (2024)
Per-Axis Weight Deltas for Frequent Model Updates
by: Kuyumdzhiev, Stefan, et al.
Published: (2025)
by: Kuyumdzhiev, Stefan, et al.
Published: (2025)
Active learning from positive and unlabeled examples
by: Mansouri, Farnam, et al.
Published: (2026)
by: Mansouri, Farnam, et al.
Published: (2026)
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
by: Dorner, Florian E., et al.
Published: (2024)
by: Dorner, Florian E., et al.
Published: (2024)
Similar Items
-
On the Interplay Between Stepsize Tuning and Progressive Sharpening
by: Roulet, Vincent, et al.
Published: (2023) -
High dimensional theory of two-phase optimizers
by: Agarwala, Atish
Published: (2026) -
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024) -
Stepping on the Edge: Curvature Aware Learning Rate Tuners
by: Roulet, Vincent, et al.
Published: (2024) -
How far away are truly hyperparameter-free learning algorithms?
by: Kasimbeg, Priya, et al.
Published: (2025)