Low-rank bias, weight decay, and model merging in neural networks
Fuente:
arXiv
Saved in:
| Main Authors: | Kuzborskij, Ilja, Yadkori, Yasin Abbasi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pointwise confidence estimation in the non-linear $\ell^2$-regularized least squares
by: Kuzborskij, Ilja, et al.
Published: (2025)
by: Kuzborskij, Ilja, et al.
Published: (2025)
To Believe or Not to Believe Your LLM
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
Best of both worlds: Stochastic & adversarial best-arm identification
by: Abbasi-Yadkori, Yasin, et al.
Published: (2026)
by: Abbasi-Yadkori, Yasin, et al.
Published: (2026)
Mitigating LLM Hallucinations via Conformal Abstention
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
Better-than-KL PAC-Bayes Bounds
by: Kuzborskij, Ilja, et al.
Published: (2024)
by: Kuzborskij, Ilja, et al.
Published: (2024)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
Sufficient Conditions for Stability of Minimum-Norm Interpolating Deep ReLU Networks
by: Harzli, Ouns El, et al.
Published: (2026)
by: Harzli, Ouns El, et al.
Published: (2026)
How does the optimizer implicitly bias the model merging loss landscape?
by: Zhang, Chenxiang, et al.
Published: (2025)
by: Zhang, Chenxiang, et al.
Published: (2025)
On-Average Stability of Multipass Preconditioned SGD and Effective Dimension
by: Vary, Simon, et al.
Published: (2026)
by: Vary, Simon, et al.
Published: (2026)
Hierarchical Reasoning Model
by: Wang, Guan, et al.
Published: (2025)
by: Wang, Guan, et al.
Published: (2025)
Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers
by: Chertkov, Andrei, et al.
Published: (2025)
by: Chertkov, Andrei, et al.
Published: (2025)
Enhancing convolutional neural network generalizability via low-rank weight approximation
by: Gao, Chenyin, et al.
Published: (2022)
by: Gao, Chenyin, et al.
Published: (2022)
Understanding the dynamics of the frequency bias in neural networks
by: Molina, Juan, et al.
Published: (2024)
by: Molina, Juan, et al.
Published: (2024)
Self-adaptive weights based on balanced residual decay rate for physics-informed neural networks and deep operator networks
by: Chen, Wenqian, et al.
Published: (2024)
by: Chen, Wenqian, et al.
Published: (2024)
On weight and variance uncertainty in neural networks for regression tasks
by: Monemi, Moein, et al.
Published: (2025)
by: Monemi, Moein, et al.
Published: (2025)
Inferring stochastic low-rank recurrent neural networks from neural data
by: Pals, Matthijs, et al.
Published: (2024)
by: Pals, Matthijs, et al.
Published: (2024)
On the effects of biased quantum random numbers on the initialization of artificial neural networks
by: Heese, Raoul, et al.
Published: (2021)
by: Heese, Raoul, et al.
Published: (2021)
A comparative analysis of a neural network with calculated weights and a neural network with random generation of weights based on the training dataset size
by: Geidarov, Polad
Published: (2025)
by: Geidarov, Polad
Published: (2025)
Emergent weight morphologies in deep neural networks
by: de Jong, Pascal, et al.
Published: (2025)
by: de Jong, Pascal, et al.
Published: (2025)
Multimodal signal fusion for stress detection using deep neural networks: a novel approach for converting 1D signals to unified 2D images
by: Hasanpoor, Yasin, et al.
Published: (2025)
by: Hasanpoor, Yasin, et al.
Published: (2025)
Posterior concentrations of fully-connected Bayesian neural networks with general priors on the weights
by: Kong, Insung, et al.
Published: (2024)
by: Kong, Insung, et al.
Published: (2024)
Analog Bayesian neural networks are insensitive to the shape of the weight distribution
by: Patel, Ravi G., et al.
Published: (2025)
by: Patel, Ravi G., et al.
Published: (2025)
Do deep neural networks utilize the weight space efficiently?
by: Koyun, Onur Can, et al.
Published: (2024)
by: Koyun, Onur Can, et al.
Published: (2024)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
by: Abbe, Emmanuel, et al.
Published: (2022)
by: Abbe, Emmanuel, et al.
Published: (2022)
Sobolev neural network with residual weighting as a surrogate in linear and non-linear mechanics
by: Kilicsoy, A. O. M., et al.
Published: (2024)
by: Kilicsoy, A. O. M., et al.
Published: (2024)
Improved weight initialization for deep and narrow feedforward neural network
by: Lee, Hyunwoo, et al.
Published: (2023)
by: Lee, Hyunwoo, et al.
Published: (2023)
Posterior and variational inference for deep neural networks with heavy-tailed weights
by: Castillo, Ismaël, et al.
Published: (2024)
by: Castillo, Ismaël, et al.
Published: (2024)
An axiomatized PDE model of deep neural networks
by: Wang, Tangjun, et al.
Published: (2023)
by: Wang, Tangjun, et al.
Published: (2023)
Weight decay induces low-rank attention layers
by: Kobayashi, Seijin, et al.
Published: (2024)
by: Kobayashi, Seijin, et al.
Published: (2024)
Multi-frequency wavefield solutions for variable velocity models using meta-learning enhanced low-rank physics-informed neural network
by: Cheng, Shijun, et al.
Published: (2025)
by: Cheng, Shijun, et al.
Published: (2025)
Convolution-weighting method for the physics-informed neural network: A Primal-Dual Optimization Perspective
by: Si, Chenhao, et al.
Published: (2025)
by: Si, Chenhao, et al.
Published: (2025)
Only relative ranks matter in weight-clustered large language models
by: Aizpurua, Borja, et al.
Published: (2026)
by: Aizpurua, Borja, et al.
Published: (2026)
Self-adaptive weighting and sampling for physics-informed neural networks
by: Chen, Wenqian, et al.
Published: (2025)
by: Chen, Wenqian, et al.
Published: (2025)
Discovering uncertainty: Gaussian constitutive neural networks with correlated weights
by: McCulloch, Jeremy A., et al.
Published: (2025)
by: McCulloch, Jeremy A., et al.
Published: (2025)
Incorporating graph neural network into route choice model
by: Ma, Yuxun, et al.
Published: (2025)
by: Ma, Yuxun, et al.
Published: (2025)
Uncertainty propagation in feed-forward neural network models
by: Diamzon, Jeremy, et al.
Published: (2025)
by: Diamzon, Jeremy, et al.
Published: (2025)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
DivMerge: A divergence-based model merging method for multi-tasking
by: Touayouch, Brahim, et al.
Published: (2025)
by: Touayouch, Brahim, et al.
Published: (2025)
Mathematical analysis of one-layer neural network with fixed biases, a new activation function and other observations
by: Macià, Fabricio, et al.
Published: (2026)
by: Macià, Fabricio, et al.
Published: (2026)
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
by: D'Amico, Francesco, et al.
Published: (2025)
by: D'Amico, Francesco, et al.
Published: (2025)
Similar Items
-
Pointwise confidence estimation in the non-linear $\ell^2$-regularized least squares
by: Kuzborskij, Ilja, et al.
Published: (2025) -
To Believe or Not to Believe Your LLM
by: Yadkori, Yasin Abbasi, et al.
Published: (2024) -
Best of both worlds: Stochastic & adversarial best-arm identification
by: Abbasi-Yadkori, Yasin, et al.
Published: (2026) -
Mitigating LLM Hallucinations via Conformal Abstention
by: Yadkori, Yasin Abbasi, et al.
Published: (2024) -
Better-than-KL PAC-Bayes Bounds
by: Kuzborskij, Ilja, et al.
Published: (2024)