Step by Step: Adaptive Gradient Descent for Training L-Lipschitz Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Sung, Kyle, Khalil, Kholood, Forman, Noah, Samu, Steven, Kratsios, Anastasis |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structure-Preserving Reconstruction of Convex Lipschitz Functionals on Hilbert Spaces from Finite Samples
by: Kratsios, Anastasis
Published: (2026)
by: Kratsios, Anastasis
Published: (2026)
Polynomial Scaling is Possible For Neural Operator Approximations of Structured Families of BSDEs
by: Furuya, Takashi, et al.
Published: (2024)
by: Furuya, Takashi, et al.
Published: (2024)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024)
by: Abuduweili, Abulikemu, et al.
Published: (2024)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023)
by: Köhne, Frederik, et al.
Published: (2023)
Is In-Context Universality Enough? MLPs are Also Universal In-Context
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Incremental Generation is Necessary and Sufficient for Universality in Flow-Based Modelling
by: Rouhvarzi, Hossein, et al.
Published: (2025)
by: Rouhvarzi, Hossein, et al.
Published: (2025)
Generative Neural Operators of Log-Complexity Can Simultaneously Solve Infinitely Many Convex Programs
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Beyond Universal Approximation Theorems: Algorithmic Uniform Approximation by Neural Networks Trained with Noisy Data
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Neural Snowflakes: Universal Latent Graph Inference via Trainable Latent Geometries
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2023)
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2023)
Quantifying The Limits of AI Reasoning: Systematic Neural Network Representations of Algorithms
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Designing Universal Causal Deep Learning Models: The Case of Infinite-Dimensional Dynamical Systems from Stochastic Analysis
by: Galimberti, Luca, et al.
Published: (2022)
by: Galimberti, Luca, et al.
Published: (2022)
Bridging the Gap Between Approximation and Learning via Optimal Approximation by ReLU MLPs of Maximal Regularity
by: Hong, Ruiyang, et al.
Published: (2024)
by: Hong, Ruiyang, et al.
Published: (2024)
Scalable Message Passing Neural Networks: No Need for Attention in Large Graph Representation Learning
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2024)
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2024)
Approximation Rates in Besov Norms and Sample-Complexity of Kolmogorov-Arnold Networks with Residual Connections
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Simultaneously Solving Infinitely Many LQ Mean Field Games In Hilbert Spaces: The Power of Neural Operators
by: Firoozi, Dena, et al.
Published: (2025)
by: Firoozi, Dena, et al.
Published: (2025)
Neural Operators Can Play Dynamic Stackelberg Games
by: Alvarez, Guillermo, et al.
Published: (2024)
by: Alvarez, Guillermo, et al.
Published: (2024)
Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region
by: Liang, Shuang, et al.
Published: (2025)
by: Liang, Shuang, et al.
Published: (2025)
Provably Faster Gradient Descent via Long Steps
by: Grimmer, Benjamin
Published: (2023)
by: Grimmer, Benjamin
Published: (2023)
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
by: Moniri, Behrad, et al.
Published: (2026)
by: Moniri, Behrad, et al.
Published: (2026)
Adaptivity Under Realizability Constraints: Comparing In-Context and Agentic Learning
by: Kratsios, Anastasis, et al.
Published: (2026)
by: Kratsios, Anastasis, et al.
Published: (2026)
Certifiable Boolean Reasoning Is Universal
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
Characterizing Overfitting in Kernel Ridgeless Regression Through the Eigenspectrum
by: Cheng, Tin Sum, et al.
Published: (2024)
by: Cheng, Tin Sum, et al.
Published: (2024)
A Comprehensive Analysis on the Learning Curve in Kernel Ridge Regression
by: Cheng, Tin Sum, et al.
Published: (2024)
by: Cheng, Tin Sum, et al.
Published: (2024)
Effectiveness of Distributed Gradient Descent with Local Steps for Overparameterized Models
by: Zhu, Heng, et al.
Published: (2024)
by: Zhu, Heng, et al.
Published: (2024)
From Logistic Regression to the Perceptron Algorithm: Exploring Gradient Descent with Large Step Sizes
by: Tyurin, Alexander
Published: (2024)
by: Tyurin, Alexander
Published: (2024)
Statistical Guarantees for Reasoning Probes on Looped Boolean Circuits
by: Kratsios, Anastasis, et al.
Published: (2026)
by: Kratsios, Anastasis, et al.
Published: (2026)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
by: Meng, Si Yi, et al.
Published: (2024)
by: Meng, Si Yi, et al.
Published: (2024)
One-Step Flow Policy Mirror Descent
by: Chen, Tianyi, et al.
Published: (2025)
by: Chen, Tianyi, et al.
Published: (2025)
Approximation and Gradient Descent Training with Neural Networks
by: Welper, G.
Published: (2024)
by: Welper, G.
Published: (2024)
A Survey on Hypergraph Neural Networks: An In-Depth and Step-By-Step Guide
by: Kim, Sunwoo, et al.
Published: (2024)
by: Kim, Sunwoo, et al.
Published: (2024)
Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
by: Meng, Si Yi, et al.
Published: (2025)
by: Meng, Si Yi, et al.
Published: (2025)
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks
by: Zhang, Chenyang, et al.
Published: (2025)
by: Zhang, Chenyang, et al.
Published: (2025)
LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
by: Arabpour, Reza, et al.
Published: (2025)
by: Arabpour, Reza, et al.
Published: (2025)
Neural Network Training via Stochastic Alternating Minimization with Trainable Step Sizes
by: Yan, Chengcheng, et al.
Published: (2025)
by: Yan, Chengcheng, et al.
Published: (2025)
One model to solve them all: 2BSDE families via neural operators
by: Furuya, Takashi, et al.
Published: (2025)
by: Furuya, Takashi, et al.
Published: (2025)
On the Lipschitz Constant of Deep Networks and Double Descent
by: Gamba, Matteo, et al.
Published: (2023)
by: Gamba, Matteo, et al.
Published: (2023)
Dual Natural Gradient Descent for Scalable Training of Physics-Informed Neural Networks
by: Jnini, Anas, et al.
Published: (2025)
by: Jnini, Anas, et al.
Published: (2025)
Stochastic Gradient Descent for Two-layer Neural Networks
by: Cao, Dinghao, et al.
Published: (2024)
by: Cao, Dinghao, et al.
Published: (2024)
Variational Stochastic Gradient Descent for Deep Neural Networks
by: Chen, Haotian, et al.
Published: (2024)
by: Chen, Haotian, et al.
Published: (2024)
Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
by: Metel, Michael R.
Published: (2022)
by: Metel, Michael R.
Published: (2022)
Similar Items
-
Structure-Preserving Reconstruction of Convex Lipschitz Functionals on Hilbert Spaces from Finite Samples
by: Kratsios, Anastasis
Published: (2026) -
Polynomial Scaling is Possible For Neural Operator Approximations of Structured Families of BSDEs
by: Furuya, Takashi, et al.
Published: (2024) -
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024) -
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023) -
Is In-Context Universality Enough? MLPs are Also Universal In-Context
by: Kratsios, Anastasis, et al.
Published: (2025)