Towards Better Generalization: Weight Decay Induces Low-rank Bias for Neural Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Ke, Yi, Chugang, Yang, Haizhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data Completion for Electrical Impedance Tomography by Conditional Diffusion Models
von: Chen, Ke, et al.
Veröffentlicht: (2026)
von: Chen, Ke, et al.
Veröffentlicht: (2026)
Curse of Dimensionality in Neural Network Optimization
von: Na, Sanghoon, et al.
Veröffentlicht: (2025)
von: Na, Sanghoon, et al.
Veröffentlicht: (2025)
Efficient Kilometer-Scale Precipitation Downscaling with Conditional Wavelet Diffusion
von: Yi, Chugang, et al.
Veröffentlicht: (2025)
von: Yi, Chugang, et al.
Veröffentlicht: (2025)
Optimal Neural Network Approximation for High-Dimensional Continuous Functions
von: Maiti, Ayan, et al.
Veröffentlicht: (2024)
von: Maiti, Ayan, et al.
Veröffentlicht: (2024)
Rethinking Inductive Bias in Geographically Neural Network Weighted Regression
von: Chen, Zhenyuan
Veröffentlicht: (2025)
von: Chen, Zhenyuan
Veröffentlicht: (2025)
Neural Networks with Sparse Activation Induced by Large Bias: Tighter Analysis with Bias-Generalized NTK
von: Yang, Hongru, et al.
Veröffentlicht: (2023)
von: Yang, Hongru, et al.
Veröffentlicht: (2023)
Towards Understanding Neural Collapse: The Effects of Batch Normalization and Weight Decay
von: Pan, Leyan, et al.
Veröffentlicht: (2023)
von: Pan, Leyan, et al.
Veröffentlicht: (2023)
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
von: Buzaglo, Gon, et al.
Veröffentlicht: (2024)
von: Buzaglo, Gon, et al.
Veröffentlicht: (2024)
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
von: Galanti, Tomer, et al.
Veröffentlicht: (2022)
von: Galanti, Tomer, et al.
Veröffentlicht: (2022)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
von: Kosson, Atli, et al.
Veröffentlicht: (2023)
von: Kosson, Atli, et al.
Veröffentlicht: (2023)
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
von: Súkeník, Peter, et al.
Veröffentlicht: (2024)
von: Súkeník, Peter, et al.
Veröffentlicht: (2024)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
Cautious Weight Decay
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
Towards Reversible Model Merging For Low-rank Weights
von: Alipour, Mohammadsajad, et al.
Veröffentlicht: (2025)
von: Alipour, Mohammadsajad, et al.
Veröffentlicht: (2025)
DP-AdamW: Investigating Decoupled Weight Decay and Bias Correction in Private Deep Learning
von: Chooi, Jay, et al.
Veröffentlicht: (2025)
von: Chooi, Jay, et al.
Veröffentlicht: (2025)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
von: Bodla, Krishna Vamshi, et al.
Veröffentlicht: (2025)
You Can Have Better Graph Neural Networks by Not Training Weights at All: Finding Untrained GNNs Tickets
von: Huang, Tianjin, et al.
Veröffentlicht: (2022)
von: Huang, Tianjin, et al.
Veröffentlicht: (2022)
Deep Grokking: Would Deep Neural Networks Generalize Better?
von: Fan, Simin, et al.
Veröffentlicht: (2024)
von: Fan, Simin, et al.
Veröffentlicht: (2024)
Neural Networks for Generating Better Local Optima in Topology Optimization
von: Herrmann, Leon, et al.
Veröffentlicht: (2024)
von: Herrmann, Leon, et al.
Veröffentlicht: (2024)
Correction of Decoupled Weight Decay
von: Chou, Jason Chuan-Chih
Veröffentlicht: (2025)
von: Chou, Jason Chuan-Chih
Veröffentlicht: (2025)
Principal Prototype Analysis on Manifold for Interpretable Reinforcement Learning
von: Vamshi, Bodla Krishna, et al.
Veröffentlicht: (2026)
von: Vamshi, Bodla Krishna, et al.
Veröffentlicht: (2026)
On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization
von: Liang, Ling, et al.
Veröffentlicht: (2024)
von: Liang, Ling, et al.
Veröffentlicht: (2024)
Finite Expression Method for Solving High-Dimensional Partial Differential Equations
von: Liang, Senwei, et al.
Veröffentlicht: (2022)
von: Liang, Senwei, et al.
Veröffentlicht: (2022)
Multi-Scale Finite Expression Method for PDEs with Oscillatory Solutions on Complex Domains
von: Hardwick, Gareth, et al.
Veröffentlicht: (2025)
von: Hardwick, Gareth, et al.
Veröffentlicht: (2025)
Teasing Apart Architecture and Initial Weights as Sources of Inductive Bias in Neural Networks
von: Bencomo, Gianluca, et al.
Veröffentlicht: (2025)
von: Bencomo, Gianluca, et al.
Veröffentlicht: (2025)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
von: Xu, Conglong, et al.
Veröffentlicht: (2025)
von: Xu, Conglong, et al.
Veröffentlicht: (2025)
On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains
von: Li, Yicheng, et al.
Veröffentlicht: (2023)
von: Li, Yicheng, et al.
Veröffentlicht: (2023)
DeepWeightFlow: Re-Basined Flow Matching for Generating Neural Network Weights
von: Gupta, Saumya, et al.
Veröffentlicht: (2026)
von: Gupta, Saumya, et al.
Veröffentlicht: (2026)
Towards Better Generalization via Distributional Input Projection Network
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
FreIE: Low-Frequency Spectral Bias in Neural Networks for Time-Series Tasks
von: Sun, Jialong, et al.
Veröffentlicht: (2025)
von: Sun, Jialong, et al.
Veröffentlicht: (2025)
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
Diffusion-Based Neural Network Weights Generation
von: Soro, Bedionita, et al.
Veröffentlicht: (2024)
von: Soro, Bedionita, et al.
Veröffentlicht: (2024)
Spectral Clustering via Orthogonalization-Free Methods
von: Pang, Qiyuan, et al.
Veröffentlicht: (2023)
von: Pang, Qiyuan, et al.
Veröffentlicht: (2023)
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
Generalization and Risk Bounds for Recurrent Neural Networks
von: Cheng, Xuewei, et al.
Veröffentlicht: (2024)
von: Cheng, Xuewei, et al.
Veröffentlicht: (2024)
The Persistence of Neural Collapse Despite Low-Rank Bias
von: Garrod, Connall, et al.
Veröffentlicht: (2024)
von: Garrod, Connall, et al.
Veröffentlicht: (2024)
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
von: Kosson, Atli, et al.
Veröffentlicht: (2025)
von: Kosson, Atli, et al.
Veröffentlicht: (2025)
Text2Weight: Bridging Natural Language and Neural Network Weight Spaces
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
Orientation-Free Neural Network-Based Bias Estimation for Low-Cost Stationary Accelerometers
von: Levin, Michal, et al.
Veröffentlicht: (2025)
von: Levin, Michal, et al.
Veröffentlicht: (2025)
AltLoRA: Towards Better Gradient Approximation in Low-Rank Adaptation with Alternating Projections
von: Yu, Xin, et al.
Veröffentlicht: (2025)
von: Yu, Xin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Data Completion for Electrical Impedance Tomography by Conditional Diffusion Models
von: Chen, Ke, et al.
Veröffentlicht: (2026) -
Curse of Dimensionality in Neural Network Optimization
von: Na, Sanghoon, et al.
Veröffentlicht: (2025) -
Efficient Kilometer-Scale Precipitation Downscaling with Conditional Wavelet Diffusion
von: Yi, Chugang, et al.
Veröffentlicht: (2025) -
Optimal Neural Network Approximation for High-Dimensional Continuous Functions
von: Maiti, Ayan, et al.
Veröffentlicht: (2024) -
Rethinking Inductive Bias in Geographically Neural Network Weighted Regression
von: Chen, Zhenyuan
Veröffentlicht: (2025)