A Rod Flow Model for Adam at the Edge of Stability
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Regis, Eric, Chewi, Sinho |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Rod Flow: A Continuous-Time Model for Gradient Descent at the Edge of Stability
par: Regis, Eric, et autres
Publié: (2026)
par: Regis, Eric, et autres
Publié: (2026)
Lectures on optimization
par: Chewi, Sinho
Publié: (2026)
par: Chewi, Sinho
Publié: (2026)
Algorithms for mean-field variational inference via polyhedral optimization in the Wasserstein space
par: Jiang, Yiheng, et autres
Publié: (2023)
par: Jiang, Yiheng, et autres
Publié: (2023)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
par: Liao, Fangshuo, et autres
Publié: (2026)
par: Liao, Fangshuo, et autres
Publié: (2026)
On the Implicit Bias of Adam
par: Cattaneo, Matias D., et autres
Publié: (2023)
par: Cattaneo, Matias D., et autres
Publié: (2023)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
par: Zhang, Fangzhao, et autres
Publié: (2026)
par: Zhang, Fangzhao, et autres
Publié: (2026)
Muon Outperforms Adam in Tail-End Associative Memory Learning
par: Wang, Shuche, et autres
Publié: (2025)
par: Wang, Shuche, et autres
Publié: (2025)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
par: Baek, Beomhan, et autres
Publié: (2025)
par: Baek, Beomhan, et autres
Publié: (2025)
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
par: Cattaneo, Matias D., et autres
Publié: (2026)
par: Cattaneo, Matias D., et autres
Publié: (2026)
Stability of Primal-Dual Gradient Flow Dynamics for Multi-Block Convex Optimization Problems
par: Ozaslan, Ibrahim K., et autres
Publié: (2024)
par: Ozaslan, Ibrahim K., et autres
Publié: (2024)
Stability of Transformers under Layer Normalization
par: Kan, Kelvin, et autres
Publié: (2025)
par: Kan, Kelvin, et autres
Publié: (2025)
Muon Dynamics as a Spectral Wasserstein Flow
par: Peyré, Gabriel
Publié: (2026)
par: Peyré, Gabriel
Publié: (2026)
Understanding Optimization in Deep Learning with Central Flows
par: Cohen, Jeremy M., et autres
Publié: (2024)
par: Cohen, Jeremy M., et autres
Publié: (2024)
gridfm-datakit-v1: A Python Library for Scalable and Realistic Power Flow and Optimal Power Flow Data Generation
par: Puech, Alban, et autres
Publié: (2025)
par: Puech, Alban, et autres
Publié: (2025)
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
par: Nie, Chengyi, et autres
Publié: (2026)
par: Nie, Chengyi, et autres
Publié: (2026)
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
par: Xiao, Nachuan, et autres
Publié: (2023)
par: Xiao, Nachuan, et autres
Publié: (2023)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
par: Sheen, Heejune, et autres
Publié: (2024)
par: Sheen, Heejune, et autres
Publié: (2024)
Stabilizing reinforcement learning control: A modular framework for optimizing over all stable behavior
par: Lawrence, Nathan P., et autres
Publié: (2023)
par: Lawrence, Nathan P., et autres
Publié: (2023)
From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
par: Ameli, Mostafa, et autres
Publié: (2025)
par: Ameli, Mostafa, et autres
Publié: (2025)
AdamZ: An Enhanced Optimisation Method for Neural Network Training
par: Zaznov, Ilia, et autres
Publié: (2024)
par: Zaznov, Ilia, et autres
Publié: (2024)
Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramér Surrogate
par: C., Simo Alami, et autres
Publié: (2025)
par: C., Simo Alami, et autres
Publié: (2025)
PGLearn -- An Open-Source Learning Toolkit for Optimal Power Flow
par: Klamkin, Michael, et autres
Publié: (2025)
par: Klamkin, Michael, et autres
Publié: (2025)
The ballistic limit of the log-Sobolev constant equals the Polyak-Łojasiewicz constant
par: Chewi, Sinho, et autres
Publié: (2024)
par: Chewi, Sinho, et autres
Publié: (2024)
Differentiable Optimization for Deep Learning-Enhanced DC Approximation of AC Optimal Power Flow
par: Rosemberg, Andrew, et autres
Publié: (2025)
par: Rosemberg, Andrew, et autres
Publié: (2025)
ARO: A New Lens On Matrix Optimization For Large Models
par: Gong, Wenbo, et autres
Publié: (2026)
par: Gong, Wenbo, et autres
Publié: (2026)
A Retention-Centric Framework for Continual Learning with Guaranteed Model Developmental Safety
par: Li, Gang, et autres
Publié: (2024)
par: Li, Gang, et autres
Publié: (2024)
From Large Language Models and Optimization to Decision Optimization CoPilot: A Research Manifesto
par: Wasserkrug, Segev, et autres
Publié: (2024)
par: Wasserkrug, Segev, et autres
Publié: (2024)
Differentiable Nonlinear Model Predictive Control
par: Frey, Jonathan, et autres
Publié: (2025)
par: Frey, Jonathan, et autres
Publié: (2025)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
par: Han, X. Y., et autres
Publié: (2025)
par: Han, X. Y., et autres
Publié: (2025)
Logistics Hub Location Optimization: A K-Means and P-Median Model Hybrid Approach Using Road Network Distances
par: Rahman, Muhammad Abdul, et autres
Publié: (2023)
par: Rahman, Muhammad Abdul, et autres
Publié: (2023)
Constructing Industrial-Scale Optimization Modeling Benchmark
par: Li, Zhong, et autres
Publié: (2026)
par: Li, Zhong, et autres
Publié: (2026)
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
par: Sheng, Jiayuan, et autres
Publié: (2025)
par: Sheng, Jiayuan, et autres
Publié: (2025)
Provable Acceleration for Diffusion Models under Minimal Assumptions
par: Li, Gen, et autres
Publié: (2024)
par: Li, Gen, et autres
Publié: (2024)
Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra
par: Kevian, Darioush, et autres
Publié: (2024)
par: Kevian, Darioush, et autres
Publié: (2024)
Data-Driven Portfolio Management for Motion Pictures Industry: A New Data-Driven Optimization Methodology Using a Large Language Model as the Expert
par: Alipour-Vaezi, Mohammad, et autres
Publié: (2024)
par: Alipour-Vaezi, Mohammad, et autres
Publié: (2024)
TaskMet: Task-Driven Metric Learning for Model Learning
par: Bansal, Dishank, et autres
Publié: (2023)
par: Bansal, Dishank, et autres
Publié: (2023)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
par: Huang, Feihu, et autres
Publié: (2026)
par: Huang, Feihu, et autres
Publié: (2026)
GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
par: Zhao, Pengxiang, et autres
Publié: (2025)
par: Zhao, Pengxiang, et autres
Publié: (2025)
Reward-Directed Score-Based Diffusion Models via q-Learning
par: Gao, Xuefeng, et autres
Publié: (2024)
par: Gao, Xuefeng, et autres
Publié: (2024)
OTAD: An Optimal Transport-Induced Robust Model for Agnostic Adversarial Attack
par: Gai, Kuo, et autres
Publié: (2024)
par: Gai, Kuo, et autres
Publié: (2024)
Documents similaires
-
Rod Flow: A Continuous-Time Model for Gradient Descent at the Edge of Stability
par: Regis, Eric, et autres
Publié: (2026) -
Lectures on optimization
par: Chewi, Sinho
Publié: (2026) -
Algorithms for mean-field variational inference via polyhedral optimization in the Wasserstein space
par: Jiang, Yiheng, et autres
Publié: (2023) -
SGD at the Edge of Stability: The Stochastic Sharpness Gap
par: Liao, Fangshuo, et autres
Publié: (2026) -
On the Implicit Bias of Adam
par: Cattaneo, Matias D., et autres
Publié: (2023)