A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Shuo, Wang, Tianhao, Wu, Beining, Li, Zhiyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structured Preconditioners in Adaptive Optimization: A Unified Analysis
by: Xie, Shuo, et al.
Published: (2025)
by: Xie, Shuo, et al.
Published: (2025)
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
by: Yadav, Robin, et al.
Published: (2025)
by: Yadav, Robin, et al.
Published: (2025)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Gradient Descent, Stochastic Optimization, and Other Tales
by: Lu, Jun
Published: (2022)
by: Lu, Jun
Published: (2022)
An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
by: Crawshaw, Michael, et al.
Published: (2025)
by: Crawshaw, Michael, et al.
Published: (2025)
Non-Euclidean Broximal Point Method: A Blueprint for Geometry-Aware Optimization
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Sample-Efficient Geometry Reconstruction from Euclidean Distances using Non-Convex Optimization
by: Ghosh, Ipsita, et al.
Published: (2024)
by: Ghosh, Ipsita, et al.
Published: (2024)
A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
Riemannian Optimization for Non-convex Euclidean Distance Geometry with Global Recovery Guarantees
by: Smith, Chandler, et al.
Published: (2024)
by: Smith, Chandler, et al.
Published: (2024)
Position: Beyond Euclidean -- Foundation Models Should Embrace Non-Euclidean Geometries
by: He, Neil, et al.
Published: (2025)
by: He, Neil, et al.
Published: (2025)
Convergence of Spectral Descent for Non-smooth Optimization
by: Yang, Yixuan, et al.
Published: (2026)
by: Yang, Yixuan, et al.
Published: (2026)
Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization
by: Cai, Yuhang, et al.
Published: (2024)
by: Cai, Yuhang, et al.
Published: (2024)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
by: Li, Bingrui, et al.
Published: (2024)
by: Li, Bingrui, et al.
Published: (2024)
A Tale of Two Symmetries: Exploring the Loss Landscape of Equivariant Models
by: Xie, YuQing, et al.
Published: (2025)
by: Xie, YuQing, et al.
Published: (2025)
Euclidean Distance Matrix Completion via Asymmetric Projected Gradient Descent
by: Li, Yicheng, et al.
Published: (2025)
by: Li, Yicheng, et al.
Published: (2025)
Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
by: Mohamadi, Mohamad Amin, et al.
Published: (2025)
by: Mohamadi, Mohamad Amin, et al.
Published: (2025)
Autoformalizing Euclidean Geometry
by: Murphy, Logan, et al.
Published: (2024)
by: Murphy, Logan, et al.
Published: (2024)
A Tale of Two Problems: Multi-Task Bilevel Learning Meets Equality Constrained Multi-Objective Optimization
by: Zhang, Zhiyao, et al.
Published: (2026)
by: Zhang, Zhiyao, et al.
Published: (2026)
A Tale of Two Cities: Pessimism and Opportunism in Offline Dynamic Pricing
by: Bian, Zeyu, et al.
Published: (2024)
by: Bian, Zeyu, et al.
Published: (2024)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024)
by: Abuduweili, Abulikemu, et al.
Published: (2024)
Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency
by: Wu, Jingfeng, et al.
Published: (2024)
by: Wu, Jingfeng, et al.
Published: (2024)
Adaptive Log-Euclidean Metrics for SPD Matrix Learning
by: Chen, Ziheng, et al.
Published: (2023)
by: Chen, Ziheng, et al.
Published: (2023)
Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning
by: Zhang, Yilang, et al.
Published: (2025)
by: Zhang, Yilang, et al.
Published: (2025)
Stochastic Adaptive Gradient Descent Without Descent
by: Aujol, Jean-François, et al.
Published: (2025)
by: Aujol, Jean-François, et al.
Published: (2025)
Provable Non-Convex Euclidean Distance Matrix Completion: Geometry, Reconstruction, and Robustness
by: Smith, Chandler, et al.
Published: (2025)
by: Smith, Chandler, et al.
Published: (2025)
Lifecycle-Aware Federated Continual Learning in Mobile Autonomous Systems
by: Wu, Beining, et al.
Published: (2026)
by: Wu, Beining, et al.
Published: (2026)
Closing the Approximation Gap of Partial AUC Optimization: A Tale of Two Formulations
by: Jiang, Yangbangyan, et al.
Published: (2025)
by: Jiang, Yangbangyan, et al.
Published: (2025)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
by: Ozkara, Kaan, et al.
Published: (2024)
by: Ozkara, Kaan, et al.
Published: (2024)
Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
Non-Euclidean Spatial Graph Neural Network
by: Zhang, Zheng, et al.
Published: (2023)
by: Zhang, Zheng, et al.
Published: (2023)
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
by: Bhatnagar, Shalabh, et al.
Published: (2022)
by: Bhatnagar, Shalabh, et al.
Published: (2022)
Incremental Sequence Labeling: A Tale of Two Shifts
by: Qiu, Shengjie, et al.
Published: (2024)
by: Qiu, Shengjie, et al.
Published: (2024)
Meta-Learning with Versatile Loss Geometries for Fast Adaptation Using Mirror Descent
by: Zhang, Yilang, et al.
Published: (2023)
by: Zhang, Yilang, et al.
Published: (2023)
Constructive Approximation under Carleman's Condition, with Applications to Smoothed Analysis
by: Koehler, Frederic, et al.
Published: (2025)
by: Koehler, Frederic, et al.
Published: (2025)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
by: Kovalev, Dmitry, et al.
Published: (2025)
by: Kovalev, Dmitry, et al.
Published: (2025)
Two-Timescale Gradient Descent Ascent Algorithms for Nonconvex Minimax Optimization
by: Lin, Tianyi, et al.
Published: (2024)
by: Lin, Tianyi, et al.
Published: (2024)
A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning
by: Zhan, Qishi, et al.
Published: (2026)
by: Zhan, Qishi, et al.
Published: (2026)
Similar Items
-
Structured Preconditioners in Adaptive Optimization: A Unified Analysis
by: Xie, Shuo, et al.
Published: (2025) -
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
by: Yadav, Robin, et al.
Published: (2025) -
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024) -
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026) -
Gradient Descent, Stochastic Optimization, and Other Tales
by: Lu, Jun
Published: (2022)