When Can You Get Away with Low Memory Adam?
Fuente:
arXiv
Saved in:
| Main Authors: | Kalra, Dayal Singh, Kirchenbauer, John, Barkeshli, Maissam, Goldstein, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
by: Kalra, Dayal Singh, et al.
Published: (2024)
by: Kalra, Dayal Singh, et al.
Published: (2024)
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
by: Kalra, Dayal Singh, et al.
Published: (2023)
by: Kalra, Dayal Singh, et al.
Published: (2023)
(How) Can Transformers Predict Pseudo-Random Numbers?
by: Tao, Tao, et al.
Published: (2025)
by: Tao, Tao, et al.
Published: (2025)
Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
by: Tao, Tao, et al.
Published: (2025)
by: Tao, Tao, et al.
Published: (2025)
On the origin of neural scaling laws: from random graphs to natural language
by: Barkeshli, Maissam, et al.
Published: (2026)
by: Barkeshli, Maissam, et al.
Published: (2026)
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
Where You Place the Norm Matters: From Prejudiced to Neutral Initializations
by: Francazi, Emanuele, et al.
Published: (2025)
by: Francazi, Emanuele, et al.
Published: (2025)
Saddle Hierarchy in Dense Associative Memory
by: Thériault, Robin, et al.
Published: (2025)
by: Thériault, Robin, et al.
Published: (2025)
Analog Physical Systems Can Exhibit Double Descent
by: Dillavou, Sam, et al.
Published: (2025)
by: Dillavou, Sam, et al.
Published: (2025)
Beyond Disorder: Unveiling Cooperativeness in Multidirectional Associative Memories
by: Alessandrelli, Andrea, et al.
Published: (2025)
by: Alessandrelli, Andrea, et al.
Published: (2025)
How Feature Learning Can Improve Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Demolition and Reinforcement of Memories in Spin-Glass-like Neural Networks
by: Ventura, Enrico
Published: (2024)
by: Ventura, Enrico
Published: (2024)
Geometric Entropy and Retrieval Phase Transitions in Continuous Thermal Dense Associative Memory
by: Petrova, Tatiana, et al.
Published: (2026)
by: Petrova, Tatiana, et al.
Published: (2026)
Learning Linear Regression with Low-Rank Tasks in-Context
by: Takanami, Kaito, et al.
Published: (2025)
by: Takanami, Kaito, et al.
Published: (2025)
The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold
by: Mao, Jialin, et al.
Published: (2023)
by: Mao, Jialin, et al.
Published: (2023)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
Thermal Robustness of Retrieval in Dense Associative Memories: LSE vs LSR Kernels
by: Petrova, Tatiana
Published: (2026)
by: Petrova, Tatiana
Published: (2026)
Nonlocal Monte Carlo via Reinforcement Learning
by: Dobrynin, Dmitrii, et al.
Published: (2025)
by: Dobrynin, Dmitrii, et al.
Published: (2025)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
by: Radhakrishnan, Anil, et al.
Published: (2025)
by: Radhakrishnan, Anil, et al.
Published: (2025)
Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
by: Xu, Yizhou, et al.
Published: (2025)
by: Xu, Yizhou, et al.
Published: (2025)
Statistical physics analysis of graph neural networks: Approaching optimality in the contextual stochastic block model
by: Duranthon, O., et al.
Published: (2025)
by: Duranthon, O., et al.
Published: (2025)
Siamese Neural Network for Label-Efficient Critical Phenomena Prediction in 3D Percolation Models
by: Wang, Shanshan, et al.
Published: (2025)
by: Wang, Shanshan, et al.
Published: (2025)
Perfect reconstruction of sparse signals using nonconvexity control and one-step RSB message passing
by: Gu, Xiaosi, et al.
Published: (2025)
by: Gu, Xiaosi, et al.
Published: (2025)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
by: Ariosto, Sebastiano
Published: (2025)
by: Ariosto, Sebastiano
Published: (2025)
A solvable model of learning generative diffusion: theory and insights
by: Cui, Hugo, et al.
Published: (2025)
by: Cui, Hugo, et al.
Published: (2025)
Graph Neural Network Approach to Predicting Magnetization in Quasi-One-Dimensional Ising Systems
by: Slavin, V., et al.
Published: (2025)
by: Slavin, V., et al.
Published: (2025)
Exploring the Energy Landscape of RBMs: Reciprocal Space Insights into Bosons, Hierarchical Learning and Symmetry Breaking
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
by: Duranthon, O., et al.
Published: (2025)
by: Duranthon, O., et al.
Published: (2025)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Learning curves theory for hierarchically compositional data with power-law distributed features
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
Physical Reinforcement Learning
by: Dillavou, Sam, et al.
Published: (2025)
by: Dillavou, Sam, et al.
Published: (2025)
Supervised and Unsupervised protocols for hetero-associative neural networks
by: Alessandrelli, Andrea, et al.
Published: (2025)
by: Alessandrelli, Andrea, et al.
Published: (2025)
Emergent weight morphologies in deep neural networks
by: de Jong, Pascal, et al.
Published: (2025)
by: de Jong, Pascal, et al.
Published: (2025)
Is Grokking a Computational Glass Relaxation?
by: Zhang, Xiaotian, et al.
Published: (2025)
by: Zhang, Xiaotian, et al.
Published: (2025)
Computing frustration and near-monotonicity in deep neural networks
by: Wendin, Joel, et al.
Published: (2025)
by: Wendin, Joel, et al.
Published: (2025)
Computational Thresholds in Multi-Modal Learning via the Spiked Matrix-Tensor Model
by: Tabanelli, Hugo, et al.
Published: (2025)
by: Tabanelli, Hugo, et al.
Published: (2025)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
by: Smart, Matthew, et al.
Published: (2025)
by: Smart, Matthew, et al.
Published: (2025)
Similar Items
-
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
by: Kalra, Dayal Singh, et al.
Published: (2024) -
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
by: Kalra, Dayal Singh, et al.
Published: (2026) -
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
by: Kalra, Dayal Singh, et al.
Published: (2023) -
(How) Can Transformers Predict Pseudo-Random Numbers?
by: Tao, Tao, et al.
Published: (2025) -
Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
by: Tao, Tao, et al.
Published: (2025)