ResNets of All Shapes and Sizes: Convergence of Training Dynamics in the Large-scale Limit
Fuente:
arXiv
Saved in:
| Main Authors: | Chaintron, Louis-Pierre, Chizat, Lénaïc, Maass, Javier |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Hidden Width of Deep ResNets: Tight Error Bounds and Phase Diagram
by: Chizat, Lénaïc
Published: (2025)
by: Chizat, Lénaïc
Published: (2025)
Deep linear networks for regression are implicitly regularized towards flat minima
by: Marion, Pierre, et al.
Published: (2024)
by: Marion, Pierre, et al.
Published: (2024)
Annealed Sinkhorn for Optimal Transport: convergence, regularization path and debiasing
by: Chizat, Lénaïc
Published: (2024)
by: Chizat, Lénaïc
Published: (2024)
The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networks
by: Chizat, Lénaïc, et al.
Published: (2023)
by: Chizat, Lénaïc, et al.
Published: (2023)
Approximating Langevin Monte Carlo with ResNet-like Neural Network architectures
by: Miranda, Charles, et al.
Published: (2023)
by: Miranda, Charles, et al.
Published: (2023)
An Exponentially Converging Particle Method for the Mixed Nash Equilibrium of Continuous Games
by: Wang, Guillaume, et al.
Published: (2022)
by: Wang, Guillaume, et al.
Published: (2022)
Doubly Regularized Entropic Wasserstein Barycenters
by: Chizat, Lénaïc
Published: (2023)
by: Chizat, Lénaïc
Published: (2023)
Scaling ResNets in the Large-depth Regime
by: Marion, Pierre, et al.
Published: (2022)
by: Marion, Pierre, et al.
Published: (2022)
Quasi-continuity method for mean-field diffusions: large deviations and central limit theorem
by: Chaintron, Louis-Pierre
Published: (2024)
by: Chaintron, Louis-Pierre
Published: (2024)
Symmetries in Overparametrized Neural Networks: A Mean-Field View
by: Maass, Javier, et al.
Published: (2024)
by: Maass, Javier, et al.
Published: (2024)
Mean-Field Langevin Dynamics for Signed Measures via a Bilevel Approach
by: Wang, Guillaume, et al.
Published: (2024)
by: Wang, Guillaume, et al.
Published: (2024)
Phase Diagram of Dropout for Two-Layer Neural Networks in the Mean-Field Regime
by: Chizat, Lénaïc, et al.
Published: (2025)
by: Chizat, Lénaïc, et al.
Published: (2025)
Regularity and stability for the Gibbs conditioning principle on path space via McKean-Vlasov control
by: Chaintron, Louis-Pierre, et al.
Published: (2024)
by: Chaintron, Louis-Pierre, et al.
Published: (2024)
On Dissipativity of Cross-Entropy Loss in Training ResNets
by: Püttschneider, Jens, et al.
Published: (2024)
by: Püttschneider, Jens, et al.
Published: (2024)
Towards an Optimal Control Perspective of ResNet Training
by: Püttschneider, Jens, et al.
Published: (2025)
by: Püttschneider, Jens, et al.
Published: (2025)
Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit
by: Riedl, Konstantin, et al.
Published: (2026)
by: Riedl, Konstantin, et al.
Published: (2026)
Geodesic convexity and strengthened functional inequalities in submanifolds of Wasserstein space
by: Chaintron, Louis-Pierre, et al.
Published: (2025)
by: Chaintron, Louis-Pierre, et al.
Published: (2025)
Optimal rate of convergence in the vanishing viscosity for quadratic Hamilton-Jacobi equations
by: Chaintron, Louis-Pierre, et al.
Published: (2025)
by: Chaintron, Louis-Pierre, et al.
Published: (2025)
Progressive Feedforward Collapse of ResNet Training
by: Wang, Sicong, et al.
Published: (2024)
by: Wang, Sicong, et al.
Published: (2024)
Quantitative Convergence of Wasserstein Gradient Flows of Kernel Mean Discrepancies
by: Chizat, Lénaïc, et al.
Published: (2026)
by: Chizat, Lénaïc, et al.
Published: (2026)
ResNets Are Deeper Than You Think
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
L-Lipschitz Gershgorin ResNet Network
by: Juston, Marius F. R., et al.
Published: (2025)
by: Juston, Marius F. R., et al.
Published: (2025)
Transformative or Conservative? Conservation laws for ResNets and Transformers
by: Marcotte, Sibylle, et al.
Published: (2025)
by: Marcotte, Sibylle, et al.
Published: (2025)
Deep Learning as a Convex Paradigm of Computation: Minimizing Circuit Size with ResNets
by: Jacot, Arthur
Published: (2025)
by: Jacot, Arthur
Published: (2025)
Uniform Scaling Limits in AdamW-Trained Transformers
by: Gibson, William, et al.
Published: (2026)
by: Gibson, William, et al.
Published: (2026)
Quantitative Local Convergence of Mean-Field Stein Variational Gradient Flow
by: Chizat, Lénaïc, et al.
Published: (2026)
by: Chizat, Lénaïc, et al.
Published: (2026)
Efficient Gravitational Wave Parameter Estimation via Knowledge Distillation: A ResNet1D-IAF Approach
by: Zhu, Xihua, et al.
Published: (2024)
by: Zhu, Xihua, et al.
Published: (2024)
Convergence, Sticking and Escape: Stochastic Dynamics Near Critical Points in SGD
by: Dudukalov, Dmitry, et al.
Published: (2025)
by: Dudukalov, Dmitry, et al.
Published: (2025)
Interpreting the Residual Stream of ResNet18
by: Longon, André
Published: (2024)
by: Longon, André
Published: (2024)
Approximation theory for 1-Lipschitz ResNets
by: Murari, Davide, et al.
Published: (2025)
by: Murari, Davide, et al.
Published: (2025)
Generalization of Scaled Deep ResNets in the Mean-Field Regime
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Bridging Neural ODE and ResNet: A Formal Error Bound for Safety Verification
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
Convergent Stochastic Training of Attention and Understanding LoRA
by: Sun, Zhengkai, et al.
Published: (2026)
by: Sun, Zhengkai, et al.
Published: (2026)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
by: Tanguy, Eloi
Published: (2023)
by: Tanguy, Eloi
Published: (2023)
Gibbs principle with infinitely many constraints: optimality conditions and stability
by: Chaintron, Louis-Pierre, et al.
Published: (2024)
by: Chaintron, Louis-Pierre, et al.
Published: (2024)
Propagation of weak log-concavity along generalised heat flows via Hamilton-Jacobi equations
by: Chaintron, Louis-Pierre, et al.
Published: (2025)
by: Chaintron, Louis-Pierre, et al.
Published: (2025)
Collective Kernel EFT for Pre-activation ResNets
by: Kawase, Hidetoshi, et al.
Published: (2026)
by: Kawase, Hidetoshi, et al.
Published: (2026)
Field theory for optimal signal propagation in ResNets
by: Fischer, Kirsten, et al.
Published: (2023)
by: Fischer, Kirsten, et al.
Published: (2023)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
Langevin Monte-Carlo Provably Learns Depth Two Neural Nets at Any Size and Data
by: Kumar, Dibyakanti, et al.
Published: (2025)
by: Kumar, Dibyakanti, et al.
Published: (2025)
Similar Items
-
The Hidden Width of Deep ResNets: Tight Error Bounds and Phase Diagram
by: Chizat, Lénaïc
Published: (2025) -
Deep linear networks for regression are implicitly regularized towards flat minima
by: Marion, Pierre, et al.
Published: (2024) -
Annealed Sinkhorn for Optimal Transport: convergence, regularization path and debiasing
by: Chizat, Lénaïc
Published: (2024) -
The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networks
by: Chizat, Lénaïc, et al.
Published: (2023) -
Approximating Langevin Monte Carlo with ResNet-like Neural Network architectures
by: Miranda, Charles, et al.
Published: (2023)