Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNets
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Zixiong, Chen, Guhan, Lai, Jianfa, Li, Bohan, Tian, Songtao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Approximating Langevin Monte Carlo with ResNet-like Neural Network architectures
di: Miranda, Charles, et al.
Pubblicazione: (2023)
di: Miranda, Charles, et al.
Pubblicazione: (2023)
Divergence of Empirical Neural Tangent Kernel in Classification Problems
di: Yu, Zixiong, et al.
Pubblicazione: (2025)
di: Yu, Zixiong, et al.
Pubblicazione: (2025)
Progressive Feedforward Collapse of ResNet Training
di: Wang, Sicong, et al.
Pubblicazione: (2024)
di: Wang, Sicong, et al.
Pubblicazione: (2024)
Overparameterization of deep ResNet: zero loss and mean-field analysis
di: Ding, Zhiyan, et al.
Pubblicazione: (2021)
di: Ding, Zhiyan, et al.
Pubblicazione: (2021)
Towards a Statistical Understanding of Neural Networks: Beyond the Neural Tangent Kernel Theories
di: Zhang, Haobo, et al.
Pubblicazione: (2024)
di: Zhang, Haobo, et al.
Pubblicazione: (2024)
Extending Mean-Field Variational Inference via Entropic Regularization: Theory and Computation
di: Wu, Bohan, et al.
Pubblicazione: (2024)
di: Wu, Bohan, et al.
Pubblicazione: (2024)
Implicit Regularization Paths of Weighted Neural Representations
di: Du, Jin-Hong, et al.
Pubblicazione: (2024)
di: Du, Jin-Hong, et al.
Pubblicazione: (2024)
Revisiting Optimism and Model Complexity in the Wake of Overparameterized Machine Learning
di: Patil, Pratik, et al.
Pubblicazione: (2024)
di: Patil, Pratik, et al.
Pubblicazione: (2024)
Batches Stabilize the Minimum Norm Risk in High Dimensional Overparameterized Linear Regression
di: Ioushua, Shahar Stein, et al.
Pubblicazione: (2023)
di: Ioushua, Shahar Stein, et al.
Pubblicazione: (2023)
Generalization of Scaled Deep ResNets in the Mean-Field Regime
di: Chen, Yihang, et al.
Pubblicazione: (2024)
di: Chen, Yihang, et al.
Pubblicazione: (2024)
Collective Kernel EFT for Pre-activation ResNets
di: Kawase, Hidetoshi, et al.
Pubblicazione: (2026)
di: Kawase, Hidetoshi, et al.
Pubblicazione: (2026)
Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent
di: Karnik, Santhosh, et al.
Pubblicazione: (2024)
di: Karnik, Santhosh, et al.
Pubblicazione: (2024)
Scaling ResNets in the Large-depth Regime
di: Marion, Pierre, et al.
Pubblicazione: (2022)
di: Marion, Pierre, et al.
Pubblicazione: (2022)
ResCP: Reservoir Conformal Prediction for Time Series Forecasting
di: Neglia, Roberto, et al.
Pubblicazione: (2025)
di: Neglia, Roberto, et al.
Pubblicazione: (2025)
Frequentist Guarantees of Distributed (Non)-Bayesian Inference
di: Wu, Bohan, et al.
Pubblicazione: (2023)
di: Wu, Bohan, et al.
Pubblicazione: (2023)
Improved Scaling Laws in Linear Regression via Data Reuse
di: Lin, Licong, et al.
Pubblicazione: (2025)
di: Lin, Licong, et al.
Pubblicazione: (2025)
On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains
di: Li, Yicheng, et al.
Pubblicazione: (2023)
di: Li, Yicheng, et al.
Pubblicazione: (2023)
A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
di: Yu, Hao
Pubblicazione: (2026)
di: Yu, Hao
Pubblicazione: (2026)
Self-Regularized Learning Methods
di: Schölpple, Max, et al.
Pubblicazione: (2026)
di: Schölpple, Max, et al.
Pubblicazione: (2026)
Implicit Regularisation in Diffusion Models: An Algorithm-Dependent Generalisation Analysis
di: Farghly, Tyler, et al.
Pubblicazione: (2025)
di: Farghly, Tyler, et al.
Pubblicazione: (2025)
Iterative Reweighted Framework Based Algorithms for Sparse Linear Regression with Generalized Elastic Net Penalty
di: Ding, Yanyun, et al.
Pubblicazione: (2024)
di: Ding, Yanyun, et al.
Pubblicazione: (2024)
Higher-Order Regularization Learning on Hypergraphs
di: Weihs, Adrien, et al.
Pubblicazione: (2025)
di: Weihs, Adrien, et al.
Pubblicazione: (2025)
Training Implicit Generative Models via an Invariant Statistical Loss
di: de Frutos, José Manuel, et al.
Pubblicazione: (2024)
di: de Frutos, José Manuel, et al.
Pubblicazione: (2024)
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
di: Súkeník, Peter, et al.
Pubblicazione: (2025)
di: Súkeník, Peter, et al.
Pubblicazione: (2025)
Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon
di: Yu, Hao
Pubblicazione: (2025)
di: Yu, Hao
Pubblicazione: (2025)
Failure of uniform laws of large numbers for subdifferentials and beyond
di: Tian, Lai, et al.
Pubblicazione: (2025)
di: Tian, Lai, et al.
Pubblicazione: (2025)
Generalized Power Priors for Improved Bayesian Inference with Historical Data
di: Kimura, Masanari, et al.
Pubblicazione: (2025)
di: Kimura, Masanari, et al.
Pubblicazione: (2025)
Implicit Regularization and Generalization in Overparameterized Neural Networks
di: Johannsen, Zeran
Pubblicazione: (2026)
di: Johannsen, Zeran
Pubblicazione: (2026)
Bagged Regularized $k$-Distances for Anomaly Detection
di: Cai, Yuchao, et al.
Pubblicazione: (2023)
di: Cai, Yuchao, et al.
Pubblicazione: (2023)
Optimal Ridge Regularization for Out-of-Distribution Prediction
di: Patil, Pratik, et al.
Pubblicazione: (2024)
di: Patil, Pratik, et al.
Pubblicazione: (2024)
Precise Asymptotics of Bagging Regularized M-estimators
di: Koriyama, Takuya, et al.
Pubblicazione: (2024)
di: Koriyama, Takuya, et al.
Pubblicazione: (2024)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
di: Zhang, Bohan, et al.
Pubblicazione: (2025)
di: Zhang, Bohan, et al.
Pubblicazione: (2025)
Arithmetic-Mean $μ$P for Modern Architectures: A Unified Learning-Rate Scale for CNNs and ResNets
di: Zhang, Haosong, et al.
Pubblicazione: (2025)
di: Zhang, Haosong, et al.
Pubblicazione: (2025)
Monte Carlo inference for semiparametric Bayesian regression
di: Kowal, Daniel R., et al.
Pubblicazione: (2023)
di: Kowal, Daniel R., et al.
Pubblicazione: (2023)
Regularization can make diffusion models more efficient
di: Taheri, Mahsa, et al.
Pubblicazione: (2025)
di: Taheri, Mahsa, et al.
Pubblicazione: (2025)
Generalization of LiNGAM that allows confounding
di: Suzuki, Joe, et al.
Pubblicazione: (2024)
di: Suzuki, Joe, et al.
Pubblicazione: (2024)
Spectrally-Corrected and Regularized QDA Classifier for Spiked Covariance Model
di: Luo, Wenya, et al.
Pubblicazione: (2025)
di: Luo, Wenya, et al.
Pubblicazione: (2025)
Finite-Particle Rates for Regularized Stein Variational Gradient Descent
di: He, Ye, et al.
Pubblicazione: (2026)
di: He, Ye, et al.
Pubblicazione: (2026)
Importance Weighting Correction of Regularized Least-Squares for Target Shift
di: Gogolashvili, Davit
Pubblicazione: (2022)
di: Gogolashvili, Davit
Pubblicazione: (2022)
Dropout Regularization Versus $\ell_2$-Penalization in the Linear Model
di: Clara, Gabriel, et al.
Pubblicazione: (2023)
di: Clara, Gabriel, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Approximating Langevin Monte Carlo with ResNet-like Neural Network architectures
di: Miranda, Charles, et al.
Pubblicazione: (2023) -
Divergence of Empirical Neural Tangent Kernel in Classification Problems
di: Yu, Zixiong, et al.
Pubblicazione: (2025) -
Progressive Feedforward Collapse of ResNet Training
di: Wang, Sicong, et al.
Pubblicazione: (2024) -
Overparameterization of deep ResNet: zero loss and mean-field analysis
di: Ding, Zhiyan, et al.
Pubblicazione: (2021) -
Towards a Statistical Understanding of Neural Networks: Beyond the Neural Tangent Kernel Theories
di: Zhang, Haobo, et al.
Pubblicazione: (2024)