Refresh-Scaling the Memory of Balanced Adam
Fuente:
arXiv
Saved in:
| Main Authors: | Fernández-Hernández, Alberto, Pérez-Corral, Cristian, Mestre, Jose I., Dolz, Manuel F., Quintana-Ortí, Enrique S. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Adam Works Better with $β_1 = β_2$: The Missing Gradient Scale Invariance Principle
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
StableGrad: Backward Scale Control without Batch Normalization
by: Mestre, Jose I., et al.
Published: (2026)
by: Mestre, Jose I., et al.
Published: (2026)
$λ$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks
by: Pérez-Corral, Cristian, et al.
Published: (2026)
by: Pérez-Corral, Cristian, et al.
Published: (2026)
OUIDecay: Adaptive Layer-wise Weight Decay for CNNs Using Online Activation Patterns
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
FedOUI: OUI-Guided Client Weighting for Federated Aggregation
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
OUI as a Structural Observable: Towards an Activation-Centric View of Neural Network Training
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
When Learning Rates Go Wrong: Early Structural Signals in PPO Actor-Critic
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
by: Mestre, Jose I., et al.
Published: (2025)
by: Mestre, Jose I., et al.
Published: (2025)
Detecting Atypical Clients in Federated Learning via Representation-Level Divergence
by: Pérez-Corral, Cristian, et al.
Published: (2026)
by: Pérez-Corral, Cristian, et al.
Published: (2026)
Regime Change Hypothesis: Foundations for Decoupled Dynamics in Neural Network Training
by: Pérez-Corral, Cristian, et al.
Published: (2026)
by: Pérez-Corral, Cristian, et al.
Published: (2026)
FedSQ: Optimized Weight Averaging via Fixed Gating
by: Pérez-Corral, Cristian, et al.
Published: (2026)
by: Pérez-Corral, Cristian, et al.
Published: (2026)
OUI Need to Talk About Weight Decay: A New Perspective on Overfitting Detection
by: Fernández-Hernández, Alberto, et al.
Published: (2025)
by: Fernández-Hernández, Alberto, et al.
Published: (2025)
Sinusoidal Initialization, Time for a New Start
by: Fernández-Hernández, Alberto, et al.
Published: (2025)
by: Fernández-Hernández, Alberto, et al.
Published: (2025)
RefreshNet: Learning Multiscale Dynamics through Hierarchical Refreshing
by: Farooq, Junaid, et al.
Published: (2024)
by: Farooq, Junaid, et al.
Published: (2024)
Quantifying uncertainty in spectral clusterings: expectations for perturbed and incomplete data
by: Dölz, Jürgen, et al.
Published: (2025)
by: Dölz, Jürgen, et al.
Published: (2025)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
Uniform Scaling Limits in AdamW-Trained Transformers
by: Gibson, William, et al.
Published: (2026)
by: Gibson, William, et al.
Published: (2026)
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
by: Malviya, Pranshu, et al.
Published: (2023)
by: Malviya, Pranshu, et al.
Published: (2023)
Studying K-FAC Heuristics by Viewing Adam through a Second-Order Lens
by: Clarke, Ross M., et al.
Published: (2023)
by: Clarke, Ross M., et al.
Published: (2023)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
Q-LocalAdam: Memory-Efficient Client-Side Adaptive Optimization for Edge Federated Learning
by: Waykole, Vedant, et al.
Published: (2026)
by: Waykole, Vedant, et al.
Published: (2026)
Data-intrinsic approximation in metric spaces
by: Dölz, Jürgen, et al.
Published: (2025)
by: Dölz, Jürgen, et al.
Published: (2025)
When Can You Get Away with Low Memory Adam?
by: Kalra, Dayal Singh, et al.
Published: (2025)
by: Kalra, Dayal Singh, et al.
Published: (2025)
CaAdam: Improving Adam optimizer using connection aware methods
by: Genet, Remi, et al.
Published: (2024)
by: Genet, Remi, et al.
Published: (2024)
Beyond Lipschitz: Data-Driven Robustness via Discrete Modulus of Continuity
by: Dölz, Jürgen, et al.
Published: (2026)
by: Dölz, Jürgen, et al.
Published: (2026)
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers
by: Ranganath, Aditya
Published: (2026)
by: Ranganath, Aditya
Published: (2026)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
by: Kim, Juno, et al.
Published: (2026)
by: Kim, Juno, et al.
Published: (2026)
No More Adam: Learning Rate Scaling at Initialization is All You Need
by: Xu, Minghao, et al.
Published: (2024)
by: Xu, Minghao, et al.
Published: (2024)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Tune My Adam, Please!
by: Athanasiadis, Theodoros, et al.
Published: (2025)
by: Athanasiadis, Theodoros, et al.
Published: (2025)
In Search of Adam's Secret Sauce
by: Orvieto, Antonio, et al.
Published: (2025)
by: Orvieto, Antonio, et al.
Published: (2025)
Batch size invariant Adam
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps
by: Ellis, Benjamin, et al.
Published: (2024)
by: Ellis, Benjamin, et al.
Published: (2024)
Adam Simplified: Bias Correction Debunked
by: Laing, Sam, et al.
Published: (2025)
by: Laing, Sam, et al.
Published: (2025)
The Implicit Bias of Adam on Separable Data
by: Zhang, Chenyang, et al.
Published: (2024)
by: Zhang, Chenyang, et al.
Published: (2024)
Keeping Code-Aware LLMs Fresh: Full Refresh, In-Context Deltas, and Incremental Fine-Tuning
by: Sharma, Pradeep Kumar, et al.
Published: (2025)
by: Sharma, Pradeep Kumar, et al.
Published: (2025)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Adaptive Preconditioners Trigger Loss Spikes in Adam
by: Bai, Zhiwei, et al.
Published: (2025)
by: Bai, Zhiwei, et al.
Published: (2025)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Similar Items
-
Why Adam Works Better with $β_1 = β_2$: The Missing Gradient Scale Invariance Principle
by: Fernández-Hernández, Alberto, et al.
Published: (2026) -
StableGrad: Backward Scale Control without Batch Normalization
by: Mestre, Jose I., et al.
Published: (2026) -
$λ$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks
by: Pérez-Corral, Cristian, et al.
Published: (2026) -
OUIDecay: Adaptive Layer-wise Weight Decay for CNNs Using Online Activation Patterns
by: Fernández-Hernández, Alberto, et al.
Published: (2026) -
FedOUI: OUI-Guided Client Weighting for Federated Aggregation
by: Fernández-Hernández, Alberto, et al.
Published: (2026)