Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
Fuente:
arXiv
Saved in:
| Main Authors: | Barboni, Raphaël, Peyré, Gabriel, Vialard, François-Xavier |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
by: Barboni, Raphaël, et al.
Published: (2024)
by: Barboni, Raphaël, et al.
Published: (2024)
Training Infinitely Deep and Wide Transformers
by: Barboni, Raphaël, et al.
Published: (2026)
by: Barboni, Raphaël, et al.
Published: (2026)
Robust Sublinear Convergence Rates for Iterative Bregman Projections
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Learning time-scales in two-layers neural networks
by: Berthier, Raphaël, et al.
Published: (2023)
by: Berthier, Raphaël, et al.
Published: (2023)
Optimal and Diffusion Transports in Machine Learning
by: Peyré, Gabriel
Published: (2025)
by: Peyré, Gabriel
Published: (2025)
Optimal Transport for Machine Learners
by: Peyré, Gabriel
Published: (2025)
by: Peyré, Gabriel
Published: (2025)
Muon Dynamics as a Spectral Wasserstein Flow
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Understanding the training of infinitely deep and wide ResNets with conditional optimal transport
by: Raphaël Barboni, et al.
Published: (2025)
by: Raphaël Barboni, et al.
Published: (2025)
Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2023)
by: Marcotte, Sibylle, et al.
Published: (2023)
On the global convergence of gradient descent for wide shallow models with bounded nonlinearities
by: Petit, Romain, et al.
Published: (2026)
by: Petit, Romain, et al.
Published: (2026)
Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2024)
by: Marcotte, Sibylle, et al.
Published: (2024)
Memory capacity of two layer neural networks with smooth activations
by: Madden, Liam, et al.
Published: (2023)
by: Madden, Liam, et al.
Published: (2023)
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
by: Zhu, Libin, et al.
Published: (2023)
by: Zhu, Libin, et al.
Published: (2023)
ICNN-enhanced 2SP: Leveraging input convex neural networks for solving two-stage stochastic programming
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
Convergence of gradient flow for learning convolutional neural networks
by: Diederen, Jona-Maria, et al.
Published: (2026)
by: Diederen, Jona-Maria, et al.
Published: (2026)
Tuning the burn-in phase in training recurrent neural networks improves their performance
by: Schiller, Julian D., et al.
Published: (2026)
by: Schiller, Julian D., et al.
Published: (2026)
Path-conditioned training: a principled way to rescale ReLU neural networks
by: Lebeurrier, Arthur, et al.
Published: (2026)
by: Lebeurrier, Arthur, et al.
Published: (2026)
Convergence of two-timescale gradient descent ascent dynamics: finite-dimensional and mean-field perspectives
by: An, Jing, et al.
Published: (2025)
by: An, Jing, et al.
Published: (2025)
Robust stabilization of polytopic systems via fast and reliable neural network-based approximations
by: Fabiani, Filippo, et al.
Published: (2022)
by: Fabiani, Filippo, et al.
Published: (2022)
Flowsheet synthesis through hierarchical reinforcement learning and graph neural networks
by: Stops, Laura, et al.
Published: (2022)
by: Stops, Laura, et al.
Published: (2022)
Tightening convex relaxations of trained neural networks: a unified approach for convex and S-shaped activations
by: Carrasco, Pablo, et al.
Published: (2024)
by: Carrasco, Pablo, et al.
Published: (2024)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Mathematical analysis of one-layer neural network with fixed biases, a new activation function and other observations
by: Macià, Fabricio, et al.
Published: (2026)
by: Macià, Fabricio, et al.
Published: (2026)
Pinet: Optimizing hard-constrained neural networks with orthogonal projection layers
by: Grontas, Panagiotis D., et al.
Published: (2025)
by: Grontas, Panagiotis D., et al.
Published: (2025)
ARM-Explainer -- Explaining and improving graph neural network predictions for the maximum clique problem using node features and association rule mining
by: Sharman, Bharat, et al.
Published: (2025)
by: Sharman, Bharat, et al.
Published: (2025)
Two-timescale Extragradient for Finding Local Minimax Points
by: Chae, Jiseok, et al.
Published: (2023)
by: Chae, Jiseok, et al.
Published: (2023)
Finite-time analysis of single-timescale actor-critic
by: Chen, Xuyang, et al.
Published: (2022)
by: Chen, Xuyang, et al.
Published: (2022)
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
by: Jentzen, Arnulf, et al.
Published: (2024)
by: Jentzen, Arnulf, et al.
Published: (2024)
Mean-field neural networks-based algorithms for McKean-Vlasov control problems *
by: Pham, Huyên, et al.
Published: (2022)
by: Pham, Huyên, et al.
Published: (2022)
From Score Matching to Diffusion: A Fine-Grained Error Analysis in the Gaussian Setting
by: Hurault, Samuel, et al.
Published: (2025)
by: Hurault, Samuel, et al.
Published: (2025)
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
Size and depth of monotone neural networks: interpolation and approximation
by: Mikulincer, Dan, et al.
Published: (2022)
by: Mikulincer, Dan, et al.
Published: (2022)
Quadratic models for understanding catapult dynamics of neural networks
by: Zhu, Libin, et al.
Published: (2022)
by: Zhu, Libin, et al.
Published: (2022)
Wasserstein distributional adversarial training for deep neural networks
by: Bai, Xingjian, et al.
Published: (2025)
by: Bai, Xingjian, et al.
Published: (2025)
Reliably-stabilizing piecewise-affine neural network controllers
by: Fabiani, Filippo, et al.
Published: (2021)
by: Fabiani, Filippo, et al.
Published: (2021)
Learning to accelerate distributed ADMM using graph neural networks
by: Doerks, Henri, et al.
Published: (2025)
by: Doerks, Henri, et al.
Published: (2025)
An analysis of optimization problems involving ReLU neural networks
by: Plate, Christoph, et al.
Published: (2025)
by: Plate, Christoph, et al.
Published: (2025)
Formulations and scalability of neural network surrogates in nonlinear optimization problems
by: Parker, Robert B., et al.
Published: (2024)
by: Parker, Robert B., et al.
Published: (2024)
A constrained optimization approach to improve robustness of neural networks
by: Zhao, Shudian, et al.
Published: (2024)
by: Zhao, Shudian, et al.
Published: (2024)
Fast Large Deformation Matching with the Energy Distance Kernel
by: Boufadene, Siwan, et al.
Published: (2025)
by: Boufadene, Siwan, et al.
Published: (2025)
Similar Items
-
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
by: Barboni, Raphaël, et al.
Published: (2024) -
Training Infinitely Deep and Wide Transformers
by: Barboni, Raphaël, et al.
Published: (2026) -
Robust Sublinear Convergence Rates for Iterative Bregman Projections
by: Peyré, Gabriel
Published: (2026) -
Learning time-scales in two-layers neural networks
by: Berthier, Raphaël, et al.
Published: (2023) -
Optimal and Diffusion Transports in Machine Learning
by: Peyré, Gabriel
Published: (2025)