Saved in:
| Main Authors: | Cattaneo, Matias D., Shigida, Boris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.02132 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2026)
by: Cattaneo, Matias D., et al.
Published: (2026)
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis
by: Cattaneo, Matias D., et al.
Published: (2025)
by: Cattaneo, Matias D., et al.
Published: (2025)
Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients
by: Alvo, Matias, et al.
Published: (2026)
by: Alvo, Matias, et al.
Published: (2026)
Dynamic Memory Based Adaptive Optimization
by: Szegedy, Balázs, et al.
Published: (2024)
by: Szegedy, Balázs, et al.
Published: (2024)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
DualSchool: How Reliable are LLMs for Optimization Education?
by: Klamkin, Michael, et al.
Published: (2025)
by: Klamkin, Michael, et al.
Published: (2025)
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
by: Lin, Jianghao, et al.
Published: (2026)
by: Lin, Jianghao, et al.
Published: (2026)
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
by: Lu, Yiyang, et al.
Published: (2025)
by: Lu, Yiyang, et al.
Published: (2025)
Federated Distributionally Robust Optimization with Non-Convex Objectives: Algorithm and Analysis
by: Jiao, Yang, et al.
Published: (2023)
by: Jiao, Yang, et al.
Published: (2023)
Unsupervised Machine Learning Hybrid Approach Integrating Linear Programming in Loss Function: A Robust Optimization Technique
by: Kiruluta, Andrew, et al.
Published: (2024)
by: Kiruluta, Andrew, et al.
Published: (2024)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
The Algorithm Configuration Problem
by: Iommazzo, Gabriele, et al.
Published: (2024)
by: Iommazzo, Gabriele, et al.
Published: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization
by: Leenders, Nick, et al.
Published: (2025)
by: Leenders, Nick, et al.
Published: (2025)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
How Well Can Transformers Emulate In-context Newton's Method?
by: Giannou, Angeliki, et al.
Published: (2024)
by: Giannou, Angeliki, et al.
Published: (2024)
The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric
by: Harvey, Thomas R.
Published: (2025)
by: Harvey, Thomas R.
Published: (2025)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
by: Schotthöfer, Steffen, et al.
Published: (2024)
by: Schotthöfer, Steffen, et al.
Published: (2024)
Pinet: Optimizing hard-constrained neural networks with orthogonal projection layers
by: Grontas, Panagiotis D., et al.
Published: (2025)
by: Grontas, Panagiotis D., et al.
Published: (2025)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
The Vizier Gaussian Process Bandit Algorithm
by: Song, Xingyou, et al.
Published: (2024)
by: Song, Xingyou, et al.
Published: (2024)
Understanding Optimization in Deep Learning with Central Flows
by: Cohen, Jeremy M., et al.
Published: (2024)
by: Cohen, Jeremy M., et al.
Published: (2024)
PID Accelerated Temporal Difference Algorithms
by: Bedaywi, Mark, et al.
Published: (2024)
by: Bedaywi, Mark, et al.
Published: (2024)
How to escape sharp minima with random perturbations
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Unified Projection-Free Algorithms for Adversarial DR-Submodular Optimization
by: Pedramfar, Mohammad, et al.
Published: (2024)
by: Pedramfar, Mohammad, et al.
Published: (2024)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
by: Han, X. Y., et al.
Published: (2025)
by: Han, X. Y., et al.
Published: (2025)
Accelerating Cutting-Plane Algorithms via Reinforcement Learning Surrogates
by: Mana, Kyle, et al.
Published: (2023)
by: Mana, Kyle, et al.
Published: (2023)
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
by: Nie, Chengyi, et al.
Published: (2026)
by: Nie, Chengyi, et al.
Published: (2026)
Optimizing the Optimizer for Physics-Informed Neural Networks and Kolmogorov-Arnold Networks
by: Kiyani, Elham, et al.
Published: (2025)
by: Kiyani, Elham, et al.
Published: (2025)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Riemannian Bilevel Optimization
by: Dutta, Sanchayan, et al.
Published: (2024)
by: Dutta, Sanchayan, et al.
Published: (2024)
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
by: Amortila, Philip, et al.
Published: (2024)
by: Amortila, Philip, et al.
Published: (2024)
Convergence of Some Convex Message Passing Algorithms to a Fixed Point
by: Voracek, Vaclav, et al.
Published: (2024)
by: Voracek, Vaclav, et al.
Published: (2024)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
by: Sharifnassab, Arsalan, et al.
Published: (2024)
by: Sharifnassab, Arsalan, et al.
Published: (2024)
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
by: Hrycej, Tomas, et al.
Published: (2025)
by: Hrycej, Tomas, et al.
Published: (2025)
From Large Language Models and Optimization to Decision Optimization CoPilot: A Research Manifesto
by: Wasserkrug, Segev, et al.
Published: (2024)
by: Wasserkrug, Segev, et al.
Published: (2024)
New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework
by: Ma, Shaocong, et al.
Published: (2026)
by: Ma, Shaocong, et al.
Published: (2026)
Similar Items
-
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2026) -
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023) -
Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis
by: Cattaneo, Matias D., et al.
Published: (2025) -
Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients
by: Alvo, Matias, et al.
Published: (2026) -
Dynamic Memory Based Adaptive Optimization
by: Szegedy, Balázs, et al.
Published: (2024)