Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xie, Shuo, Mohamadi, Mohamad Amin, Li, Zhiyuan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
par: Xie, Shuo, et autres
Publié: (2024)
par: Xie, Shuo, et autres
Publié: (2024)
Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
par: Mohamadi, Mohamad Amin, et autres
Publié: (2025)
par: Mohamadi, Mohamad Amin, et autres
Publié: (2025)
Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition
par: Mohamadi, Mohamad Amin, et autres
Publié: (2024)
par: Mohamadi, Mohamad Amin, et autres
Publié: (2024)
Adaptive Preconditioners Trigger Loss Spikes in Adam
par: Bai, Zhiwei, et autres
Publié: (2025)
par: Bai, Zhiwei, et autres
Publié: (2025)
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
par: Xie, Shuo, et autres
Publié: (2025)
par: Xie, Shuo, et autres
Publié: (2025)
Asymptotic Behavior of Adversarial Training Estimator under $\ell_\infty$-Perturbation
par: Xie, Yiling, et autres
Publié: (2024)
par: Xie, Yiling, et autres
Publié: (2024)
Improved Distribution Estimation in $\ell_\infty$
par: Cohen, Doron, et autres
Publié: (2026)
par: Cohen, Doron, et autres
Publié: (2026)
Structured Preconditioners in Adaptive Optimization: A Unified Analysis
par: Xie, Shuo, et autres
Publié: (2025)
par: Xie, Shuo, et autres
Publié: (2025)
Patch-wise Structural Loss for Time Series Forecasting
par: Kudrat, Dilfira, et autres
Publié: (2025)
par: Kudrat, Dilfira, et autres
Publié: (2025)
Beyond $\ell_2$-norm and $\ell_\infty$-norm: A Curvature-Inspired $\ell_p$-Norm Scheme for Deep Neural Networks
par: Xu, Jianhao, et autres
Publié: (2026)
par: Xu, Jianhao, et autres
Publié: (2026)
Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer
par: Xie, Yan-Feng, et autres
Publié: (2026)
par: Xie, Yan-Feng, et autres
Publié: (2026)
LossLens: Diagnostics for Machine Learning through Loss Landscape Visual Analytics
par: Xie, Tiankai, et autres
Publié: (2024)
par: Xie, Tiankai, et autres
Publié: (2024)
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
par: Wen, Kaiyue, et autres
Publié: (2024)
par: Wen, Kaiyue, et autres
Publié: (2024)
Sensitivity Analysis On Loss Landscape
par: Faroz, Salman
Publié: (2024)
par: Faroz, Salman
Publié: (2024)
A Determinantal Approach to a Sharp $\ell^1-\ell^\infty-\ell^2$ Norm Inequality
par: Benitez, Jose Antonio Lara
Publié: (2026)
par: Benitez, Jose Antonio Lara
Publié: (2026)
Universally Empowering Zeroth-Order Optimization via Adaptive Layer-wise Sampling
par: Wang, Fei, et autres
Publié: (2026)
par: Wang, Fei, et autres
Publié: (2026)
Adaptive Refinement Protocols for Distributed Distribution Estimation under $\ell^p$-Losses
par: Yuan, Deheng, et autres
Publié: (2024)
par: Yuan, Deheng, et autres
Publié: (2024)
OLion: Approaching the Hadamard Ideal by Intersecting Spectral and $\ell_{\infty}$ Implicit Biases
par: Wang, Zixiao, et autres
Publié: (2026)
par: Wang, Zixiao, et autres
Publié: (2026)
Landscaper: Understanding Loss Landscapes Through Multi-Dimensional Topological Analysis
par: Chen, Jiaqing, et autres
Publié: (2026)
par: Chen, Jiaqing, et autres
Publié: (2026)
A Tale of Two Symmetries: Exploring the Loss Landscape of Equivariant Models
par: Xie, YuQing, et autres
Publié: (2025)
par: Xie, YuQing, et autres
Publié: (2025)
Curvature in the Looking-Glass: Optimal Methods to Exploit Curvature of Expectation in the Loss Landscape
par: Duersch, Jed A., et autres
Publié: (2024)
par: Duersch, Jed A., et autres
Publié: (2024)
CP Loss: Channel-wise Perceptual Loss for Time Series Forecasting
par: Zha, Yaohua, et autres
Publié: (2026)
par: Zha, Yaohua, et autres
Publié: (2026)
Visualizing Loss Functions as Topological Landscape Profiles
par: Geniesse, Caleb, et autres
Publié: (2024)
par: Geniesse, Caleb, et autres
Publié: (2024)
Visualizing, Rethinking, and Mining the Loss Landscape of Deep Neural Networks
par: Xu, Yichu, et autres
Publié: (2024)
par: Xu, Yichu, et autres
Publié: (2024)
On the Hyperparameter Loss Landscapes of Machine Learning Models: An Exploratory Study
par: Huang, Mingyu, et autres
Publié: (2023)
par: Huang, Mingyu, et autres
Publié: (2023)
GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
par: Kim, Jinuk, et autres
Publié: (2025)
par: Kim, Jinuk, et autres
Publié: (2025)
Estimating Higher-Order Mixed Memberships via the $\ell_{2,\infty}$ Tensor Perturbation Bound
par: Agterberg, Joshua, et autres
Publié: (2022)
par: Agterberg, Joshua, et autres
Publié: (2022)
Exploiting Preferences in Loss Functions for Sequential Recommendation via Weak Transitivity
par: Chung, Hyunsoo, et autres
Publié: (2024)
par: Chung, Hyunsoo, et autres
Publié: (2024)
Evaluating Loss Landscapes from a Topology Perspective
par: Xie, Tiankai, et autres
Publié: (2024)
par: Xie, Tiankai, et autres
Publié: (2024)
There is a Singularity in the Loss Landscape
par: Lowell, Mark
Publié: (2022)
par: Lowell, Mark
Publié: (2022)
On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
par: Li, Huan, et autres
Publié: (2025)
par: Li, Huan, et autres
Publié: (2025)
Stable Coresets via Posterior Sampling: Aligning Induced and Full Loss Landscapes
par: Chang, Wei-Kai, et autres
Publié: (2025)
par: Chang, Wei-Kai, et autres
Publié: (2025)
Early-Warning Signals of Grokking via Loss-Landscape Geometry
par: Xu, Yongzhong
Publié: (2026)
par: Xu, Yongzhong
Publié: (2026)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
par: Bushnaq, Lucius, et autres
Publié: (2024)
par: Bushnaq, Lucius, et autres
Publié: (2024)
Paths and Ambient Spaces in Neural Loss Landscapes
par: Dold, Daniel, et autres
Publié: (2025)
par: Dold, Daniel, et autres
Publié: (2025)
Near-Linear Time Projection onto the $\ell_{1,\infty}$ Ball; Application to Sparse Autoencoders
par: Perez, Guillaume, et autres
Publié: (2023)
par: Perez, Guillaume, et autres
Publié: (2023)
On the Computational Landscape of Replicable Learning
par: Kalavasis, Alkis, et autres
Publié: (2024)
par: Kalavasis, Alkis, et autres
Publié: (2024)
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
par: Yadav, Robin, et autres
Publié: (2025)
par: Yadav, Robin, et autres
Publié: (2025)
FedLWS: Federated Learning with Adaptive Layer-wise Weight Shrinking
par: Shi, Changlong, et autres
Publié: (2025)
par: Shi, Changlong, et autres
Publié: (2025)
On Suppressing Range of Adaptive Stepsizes of Adam to Improve Generalisation Performance
par: Zhang, Guoqiang
Publié: (2023)
par: Zhang, Guoqiang
Publié: (2023)
Documents similaires
-
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
par: Xie, Shuo, et autres
Publié: (2024) -
Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
par: Mohamadi, Mohamad Amin, et autres
Publié: (2025) -
Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition
par: Mohamadi, Mohamad Amin, et autres
Publié: (2024) -
Adaptive Preconditioners Trigger Loss Spikes in Adam
par: Bai, Zhiwei, et autres
Publié: (2025) -
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
par: Xie, Shuo, et autres
Publié: (2025)