Quadratic models for understanding catapult dynamics of neural networks
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Libin, Liu, Chaoyue, Radhakrishnan, Adityanarayanan, Belkin, Mikhail |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
by: Zhu, Libin, et al.
Published: (2023)
by: Zhu, Libin, et al.
Published: (2023)
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
by: Mallinar, Neil, et al.
Published: (2024)
by: Mallinar, Neil, et al.
Published: (2024)
Mirror Descent on Reproducing Kernel Banach Spaces
by: Kumar, Akash, et al.
Published: (2024)
by: Kumar, Akash, et al.
Published: (2024)
Linear Recursive Feature Machines provably recover low-rank matrices
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
Efficient model predictive control for nonlinear systems modelled by deep neural networks
by: Lan, Jianglin
Published: (2024)
by: Lan, Jianglin
Published: (2024)
Solving a class of stochastic optimal control problems by physics-informed neural networks
by: Jiao, Zhe, et al.
Published: (2024)
by: Jiao, Zhe, et al.
Published: (2024)
Approximate non-linear model predictive control with safety-augmented neural networks
by: Hose, Henrik, et al.
Published: (2023)
by: Hose, Henrik, et al.
Published: (2023)
Size and depth of monotone neural networks: interpolation and approximation
by: Mikulincer, Dan, et al.
Published: (2022)
by: Mikulincer, Dan, et al.
Published: (2022)
Convergence of gradient flow for learning convolutional neural networks
by: Diederen, Jona-Maria, et al.
Published: (2026)
by: Diederen, Jona-Maria, et al.
Published: (2026)
Reliably-stabilizing piecewise-affine neural network controllers
by: Fabiani, Filippo, et al.
Published: (2021)
by: Fabiani, Filippo, et al.
Published: (2021)
Context-Scaling versus Task-Scaling in In-Context Learning
by: Abedsoltan, Amirhesam, et al.
Published: (2024)
by: Abedsoltan, Amirhesam, et al.
Published: (2024)
Formulations and scalability of neural network surrogates in nonlinear optimization problems
by: Parker, Robert B., et al.
Published: (2024)
by: Parker, Robert B., et al.
Published: (2024)
Learning to accelerate distributed ADMM using graph neural networks
by: Doerks, Henri, et al.
Published: (2025)
by: Doerks, Henri, et al.
Published: (2025)
A constrained optimization approach to improve robustness of neural networks
by: Zhao, Shudian, et al.
Published: (2024)
by: Zhao, Shudian, et al.
Published: (2024)
An analysis of optimization problems involving ReLU neural networks
by: Plate, Christoph, et al.
Published: (2025)
by: Plate, Christoph, et al.
Published: (2025)
Approximation and interpolation of deep neural networks
by: Constantinescu, Vlad-Raul, et al.
Published: (2023)
by: Constantinescu, Vlad-Raul, et al.
Published: (2023)
ICNN-enhanced 2SP: Leveraging input convex neural networks for solving two-stage stochastic programming
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
Flowsheet synthesis through hierarchical reinforcement learning and graph neural networks
by: Stops, Laura, et al.
Published: (2022)
by: Stops, Laura, et al.
Published: (2022)
The duality structure gradient descent algorithm: analysis and applications to neural networks
by: Flynn, Thomas
Published: (2017)
by: Flynn, Thomas
Published: (2017)
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
by: Zhang, Haihan, et al.
Published: (2024)
by: Zhang, Haihan, et al.
Published: (2024)
A neural network-based approach to hybrid systems identification for control
by: Fabiani, Filippo, et al.
Published: (2024)
by: Fabiani, Filippo, et al.
Published: (2024)
Tuning the burn-in phase in training recurrent neural networks improves their performance
by: Schiller, Julian D., et al.
Published: (2026)
by: Schiller, Julian D., et al.
Published: (2026)
MIQCQP reformulation of the ReLU neural networks Lipschitz constant estimation problem
by: Sbihi, Mohammed, et al.
Published: (2024)
by: Sbihi, Mohammed, et al.
Published: (2024)
Verifying message-passing neural networks via topology-based bounds tightening
by: Hojny, Christopher, et al.
Published: (2024)
by: Hojny, Christopher, et al.
Published: (2024)
Convergence of continuous-time stochastic gradient descent with applications to deep neural networks
by: Lugosi, Gabor, et al.
Published: (2024)
by: Lugosi, Gabor, et al.
Published: (2024)
Path-conditioned training: a principled way to rescale ReLU neural networks
by: Lebeurrier, Arthur, et al.
Published: (2026)
by: Lebeurrier, Arthur, et al.
Published: (2026)
Expressivity of Quadratic Neural ODEs
by: Hanson, Joshua, et al.
Published: (2025)
by: Hanson, Joshua, et al.
Published: (2025)
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
Smart energy management: process structure-based hybrid neural networks for optimal scheduling and economic predictive control in integrated systems
by: Wu, Long, et al.
Published: (2024)
by: Wu, Long, et al.
Published: (2024)
Robust stabilization of polytopic systems via fast and reliable neural network-based approximations
by: Fabiani, Filippo, et al.
Published: (2022)
by: Fabiani, Filippo, et al.
Published: (2022)
Iteratively reweighted kernel machines efficiently learn sparse functions
by: Zhu, Libin, et al.
Published: (2025)
by: Zhu, Libin, et al.
Published: (2025)
Insights on Muon from Simple Quadratics
by: Gonon, Antoine, et al.
Published: (2026)
by: Gonon, Antoine, et al.
Published: (2026)
A Penalty Approach for Differentiation Through Black-Box Quadratic Programming Solvers
by: Linghu, Yuxuan, et al.
Published: (2026)
by: Linghu, Yuxuan, et al.
Published: (2026)
Expressive Power of Graph Neural Networks for (Mixed-Integer) Quadratic Programs
by: Chen, Ziang, et al.
Published: (2024)
by: Chen, Ziang, et al.
Published: (2024)
Improved Physics-informed neural networks loss function regularization with a variance-based term
by: Hanna, John M., et al.
Published: (2024)
by: Hanna, John M., et al.
Published: (2024)
Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
by: Barboni, Raphaël, et al.
Published: (2025)
by: Barboni, Raphaël, et al.
Published: (2025)
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
by: An, Jing, et al.
Published: (2023)
by: An, Jing, et al.
Published: (2023)
Tightening convex relaxations of trained neural networks: a unified approach for convex and S-shaped activations
by: Carrasco, Pablo, et al.
Published: (2024)
by: Carrasco, Pablo, et al.
Published: (2024)
On bounds for norms of reparameterized ReLU artificial neural network parameters: sums of fractional powers of the Lipschitz norm control the network parameter vector
by: Jentzen, Arnulf, et al.
Published: (2022)
by: Jentzen, Arnulf, et al.
Published: (2022)
Similar Items
-
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
by: Zhu, Libin, et al.
Published: (2023) -
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
by: Mallinar, Neil, et al.
Published: (2024) -
Mirror Descent on Reproducing Kernel Banach Spaces
by: Kumar, Akash, et al.
Published: (2024) -
Linear Recursive Feature Machines provably recover low-rank matrices
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024) -
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
by: Beaglehole, Daniel, et al.
Published: (2025)