Penalising the biases in norm regularisation enforces sparsity
Fuente:
arXiv
Saved in:
| Main Authors: | Boursier, Etienne, Flammarion, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simplicity bias and optimization threshold in two-layer ReLU networks
by: Boursier, Etienne, et al.
Published: (2024)
by: Boursier, Etienne, et al.
Published: (2024)
Early alignment in two-layer networks training is a two-edged sword
by: Boursier, Etienne, et al.
Published: (2024)
by: Boursier, Etienne, et al.
Published: (2024)
First-order ANIL provably learns representations despite overparametrization
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
A survey on multi-player bandits
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
Transformers Learn Latent Mixture Models In-Context via Mirror Descent
by: D'Angelo, Francesco, et al.
Published: (2026)
by: D'Angelo, Francesco, et al.
Published: (2026)
Benignity of loss landscape with weight decay requires both large overparametrization and initialization
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points
by: Varre, Aditya, et al.
Published: (2025)
by: Varre, Aditya, et al.
Published: (2025)
Learning Parametric Distributions from Samples and Preferences
by: Jourdan, Marc, et al.
Published: (2025)
by: Jourdan, Marc, et al.
Published: (2025)
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024)
by: Papazov, Hristo, et al.
Published: (2024)
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
by: Varre, Aditya, et al.
Published: (2026)
by: Varre, Aditya, et al.
Published: (2026)
(How) Learning Rates Regulate Catastrophic Overtraining
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias
by: Town, James, et al.
Published: (2026)
by: Town, James, et al.
Published: (2026)
Selective Induction Heads: How Transformers Select Causal Structures In Context
by: D'Angelo, Francesco, et al.
Published: (2025)
by: D'Angelo, Francesco, et al.
Published: (2025)
Learning Algorithms in the Limit
by: Papazov, Hristo, et al.
Published: (2025)
by: Papazov, Hristo, et al.
Published: (2025)
Approximate information maximization for bandit games
by: Barbier-Chebbah, Alex, et al.
Published: (2023)
by: Barbier-Chebbah, Alex, et al.
Published: (2023)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Exact Learning of Arithmetic with Differentiable Agents
by: Papazov, Hristo, et al.
Published: (2025)
by: Papazov, Hristo, et al.
Published: (2025)
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
Why Do We Need Weight Decay in Modern Deep Learning?
by: D'Angelo, Francesco, et al.
Published: (2023)
by: D'Angelo, Francesco, et al.
Published: (2023)
Inference for relative sparsity
by: Weisenthal, Samuel J., et al.
Published: (2023)
by: Weisenthal, Samuel J., et al.
Published: (2023)
Effective regions and kernels in continuous sparse regularisation, with application to sketched mixtures
by: De Castro, Yohann, et al.
Published: (2025)
by: De Castro, Yohann, et al.
Published: (2025)
Incremental Learning of Sparse Attention Patterns in Transformers
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
Long-Context Linear System Identification
by: Yüksel, Oğuz Kaan, et al.
Published: (2024)
by: Yüksel, Oğuz Kaan, et al.
Published: (2024)
Online Decision-Focused Learning
by: Capitaine, Aymeric, et al.
Published: (2025)
by: Capitaine, Aymeric, et al.
Published: (2025)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025)
by: Schlarmann, Christian, et al.
Published: (2025)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
by: Neuhaus, Yannic, et al.
Published: (2026)
by: Neuhaus, Yannic, et al.
Published: (2026)
Optimal Design for Reward Modeling in RLHF
by: Scheid, Antoine, et al.
Published: (2024)
by: Scheid, Antoine, et al.
Published: (2024)
Learning to Mitigate Externalities: the Coase Theorem with Hindsight Rationality
by: Scheid, Antoine, et al.
Published: (2024)
by: Scheid, Antoine, et al.
Published: (2024)
Can sparsity improve the privacy of neural networks?
by: Gonon, Antoine, et al.
Published: (2023)
by: Gonon, Antoine, et al.
Published: (2023)
Unconditional flow-based time series generation with equivariance-regularised latent spaces
by: Reyes, Camilo Carvajal, et al.
Published: (2026)
by: Reyes, Camilo Carvajal, et al.
Published: (2026)
Group selection and shrinkage: Structured sparsity for semiparametric additive models
by: Thompson, Ryan, et al.
Published: (2021)
by: Thompson, Ryan, et al.
Published: (2021)
On the necessity of adaptive regularisation:Optimal anytime online learning on $\boldsymbol{\ell_p}$-balls
by: Johnson, Emmeran, et al.
Published: (2025)
by: Johnson, Emmeran, et al.
Published: (2025)
Manifold-regularised Large-Margin $\ell_p$-SVDD for Multidimensional Time Series Anomaly Detection
by: Arashloo, Shervin Rahimzadeh
Published: (2025)
by: Arashloo, Shervin Rahimzadeh
Published: (2025)
On sparsity, extremal structure, and monotonicity properties of Wasserstein and Gromov-Wasserstein optimal transport plans
by: Vayer, Titouan
Published: (2026)
by: Vayer, Titouan
Published: (2026)
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
by: Ghazanfari, Sara, et al.
Published: (2024)
by: Ghazanfari, Sara, et al.
Published: (2024)
Similar Items
-
Simplicity bias and optimization threshold in two-layer ReLU networks
by: Boursier, Etienne, et al.
Published: (2024) -
Early alignment in two-layer networks training is a two-edged sword
by: Boursier, Etienne, et al.
Published: (2024) -
First-order ANIL provably learns representations despite overparametrization
by: Yüksel, Oğuz Kaan, et al.
Published: (2023) -
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
by: Boursier, Etienne, et al.
Published: (2022) -
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
by: Boursier, Etienne, et al.
Published: (2025)