Transformative or Conservative? Conservation laws for ResNets and Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Marcotte, Sibylle, Gribonval, Rémi, Peyré, Gabriel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2023)
by: Marcotte, Sibylle, et al.
Published: (2023)
Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2024)
by: Marcotte, Sibylle, et al.
Published: (2024)
Intrinsic training dynamics of deep neural networks
by: Marcotte, Sibylle, et al.
Published: (2025)
by: Marcotte, Sibylle, et al.
Published: (2025)
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
by: Barboni, Raphaël, et al.
Published: (2024)
by: Barboni, Raphaël, et al.
Published: (2024)
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
by: Súkeník, Peter, et al.
Published: (2025)
by: Súkeník, Peter, et al.
Published: (2025)
Scaling ResNets in the Large-depth Regime
by: Marion, Pierre, et al.
Published: (2022)
by: Marion, Pierre, et al.
Published: (2022)
SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks
by: Shi, Xinyu, et al.
Published: (2024)
by: Shi, Xinyu, et al.
Published: (2024)
ResNets Are Deeper Than You Think
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
L-Lipschitz Gershgorin ResNet Network
by: Juston, Marius F. R., et al.
Published: (2025)
by: Juston, Marius F. R., et al.
Published: (2025)
Approximation theory for 1-Lipschitz ResNets
by: Murari, Davide, et al.
Published: (2025)
by: Murari, Davide, et al.
Published: (2025)
Interpreting the Residual Stream of ResNet18
by: Longon, André
Published: (2024)
by: Longon, André
Published: (2024)
Revisiting RIP guarantees for sketching operators on mixture models
by: Belhadji, Ayoub, et al.
Published: (2023)
by: Belhadji, Ayoub, et al.
Published: (2023)
Sketch and shift: a robust decoder for compressive clustering
by: Belhadji, Ayoub, et al.
Published: (2023)
by: Belhadji, Ayoub, et al.
Published: (2023)
Generalization of Scaled Deep ResNets in the Mean-Field Regime
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Towards Understanding the Universality of Transformers for Next-Token Prediction
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
A Comparative Study of CNN, ResNet, and Vision Transformers for Multi-Classification of Chest Diseases
by: Jain, Ananya, et al.
Published: (2024)
by: Jain, Ananya, et al.
Published: (2024)
Towards an Optimal Control Perspective of ResNet Training
by: Püttschneider, Jens, et al.
Published: (2025)
by: Püttschneider, Jens, et al.
Published: (2025)
On Dissipativity of Cross-Entropy Loss in Training ResNets
by: Püttschneider, Jens, et al.
Published: (2024)
by: Püttschneider, Jens, et al.
Published: (2024)
Field theory for optimal signal propagation in ResNets
by: Fischer, Kirsten, et al.
Published: (2023)
by: Fischer, Kirsten, et al.
Published: (2023)
Collective Kernel EFT for Pre-activation ResNets
by: Kawase, Hidetoshi, et al.
Published: (2026)
by: Kawase, Hidetoshi, et al.
Published: (2026)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
Hamiltonian Mechanics of Feature Learning: Bottleneck Structure in Leaky ResNets
by: Jacot, Arthur, et al.
Published: (2024)
by: Jacot, Arthur, et al.
Published: (2024)
Overparameterization of deep ResNet: zero loss and mean-field analysis
by: Ding, Zhiyan, et al.
Published: (2021)
by: Ding, Zhiyan, et al.
Published: (2021)
Understanding the training of infinitely deep and wide ResNets with conditional optimal transport
by: Raphaël Barboni, et al.
Published: (2025)
by: Raphaël Barboni, et al.
Published: (2025)
Poly-MgNet: Polynomial Building Blocks in Multigrid-Inspired ResNets
by: van Betteray, Antonia, et al.
Published: (2025)
by: van Betteray, Antonia, et al.
Published: (2025)
Progressive Feedforward Collapse of ResNet Training
by: Wang, Sicong, et al.
Published: (2024)
by: Wang, Sicong, et al.
Published: (2024)
BrainRotViT: Transformer-ResNet Hybrid for Explainable Modeling of Brain Aging from 3D sMRI
by: Jalal, Wasif, et al.
Published: (2025)
by: Jalal, Wasif, et al.
Published: (2025)
Transformers are Universal In-context Learners
by: Furuya, Takashi, et al.
Published: (2024)
by: Furuya, Takashi, et al.
Published: (2024)
Naturally Computed Scale Invariance in the Residual Stream of ResNet18
by: Longon, André
Published: (2025)
by: Longon, André
Published: (2025)
Layer-Wise Relevance Propagation with Conservation Property for ResNet
by: Otsuki, Seitaro, et al.
Published: (2024)
by: Otsuki, Seitaro, et al.
Published: (2024)
ResNets of All Shapes and Sizes: Convergence of Training Dynamics in the Large-scale Limit
by: Chaintron, Louis-Pierre, et al.
Published: (2026)
by: Chaintron, Louis-Pierre, et al.
Published: (2026)
Approximating Langevin Monte Carlo with ResNet-like Neural Network architectures
by: Miranda, Charles, et al.
Published: (2023)
by: Miranda, Charles, et al.
Published: (2023)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
ActNetFormer: Transformer-ResNet Hybrid Method for Semi-Supervised Action Recognition in Videos
by: Dass, Sharana Dharshikgan Suresh, et al.
Published: (2024)
by: Dass, Sharana Dharshikgan Suresh, et al.
Published: (2024)
Bridging Neural ODE and ResNet: A Formal Error Bound for Safety Verification
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
Deep Learning as a Convex Paradigm of Computation: Minimizing Circuit Size with ResNets
by: Jacot, Arthur
Published: (2025)
by: Jacot, Arthur
Published: (2025)
Invertible ResNets for Inverse Imaging Problems: Competitive Performance with Provable Regularization Properties
by: Arndt, Clemens, et al.
Published: (2024)
by: Arndt, Clemens, et al.
Published: (2024)
Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNets
by: Yu, Zixiong, et al.
Published: (2024)
by: Yu, Zixiong, et al.
Published: (2024)
Cardioformer: Advancing AI in ECG Analysis with Multi-Granularity Patching and ResNet
by: Mobin, Md Kamrujjaman, et al.
Published: (2025)
by: Mobin, Md Kamrujjaman, et al.
Published: (2025)
Path-conditioned training: a principled way to rescale ReLU neural networks
by: Lebeurrier, Arthur, et al.
Published: (2026)
by: Lebeurrier, Arthur, et al.
Published: (2026)
Similar Items
-
Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2023) -
Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2024) -
Intrinsic training dynamics of deep neural networks
by: Marcotte, Sibylle, et al.
Published: (2025) -
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
by: Barboni, Raphaël, et al.
Published: (2024) -
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
by: Súkeník, Peter, et al.
Published: (2025)