Early alignment in two-layer networks training is a two-edged sword
Fuente:
arXiv
Saved in:
| Main Authors: | Boursier, Etienne, Flammarion, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simplicity bias and optimization threshold in two-layer ReLU networks
by: Boursier, Etienne, et al.
Published: (2024)
by: Boursier, Etienne, et al.
Published: (2024)
Penalising the biases in norm regularisation enforces sparsity
by: Boursier, Etienne, et al.
Published: (2023)
by: Boursier, Etienne, et al.
Published: (2023)
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
First-order ANIL provably learns representations despite overparametrization
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
A survey on multi-player bandits
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
by: Barboni, Raphaël, et al.
Published: (2025)
by: Barboni, Raphaël, et al.
Published: (2025)
Transformers Learn Latent Mixture Models In-Context via Mirror Descent
by: D'Angelo, Francesco, et al.
Published: (2026)
by: D'Angelo, Francesco, et al.
Published: (2026)
Benignity of loss landscape with weight decay requires both large overparametrization and initialization
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025)
by: Schlarmann, Christian, et al.
Published: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points
by: Varre, Aditya, et al.
Published: (2025)
by: Varre, Aditya, et al.
Published: (2025)
Learning Parametric Distributions from Samples and Preferences
by: Jourdan, Marc, et al.
Published: (2025)
by: Jourdan, Marc, et al.
Published: (2025)
Generalization error bounds for two-layer neural networks with Lipschitz loss function
by: Nguwi, Jiang Yu, et al.
Published: (2026)
by: Nguwi, Jiang Yu, et al.
Published: (2026)
Generalization error property of infoGAN for two-layer neural network
by: Hasan, Mahmud, et al.
Published: (2023)
by: Hasan, Mahmud, et al.
Published: (2023)
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024)
by: Papazov, Hristo, et al.
Published: (2024)
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
by: Varre, Aditya, et al.
Published: (2026)
by: Varre, Aditya, et al.
Published: (2026)
(How) Learning Rates Regulate Catastrophic Overtraining
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias
by: Town, James, et al.
Published: (2026)
by: Town, James, et al.
Published: (2026)
Selective Induction Heads: How Transformers Select Causal Structures In Context
by: D'Angelo, Francesco, et al.
Published: (2025)
by: D'Angelo, Francesco, et al.
Published: (2025)
Explicit integral representations and quantitative bounds for two-layer ReLU networks
by: Lee, Anthony
Published: (2026)
by: Lee, Anthony
Published: (2026)
A duality framework for analyzing random feature and two-layer neural networks
by: Chen, Hongrui, et al.
Published: (2023)
by: Chen, Hongrui, et al.
Published: (2023)
Learning Algorithms in the Limit
by: Papazov, Hristo, et al.
Published: (2025)
by: Papazov, Hristo, et al.
Published: (2025)
Approximate information maximization for bandit games
by: Barbier-Chebbah, Alex, et al.
Published: (2023)
by: Barbier-Chebbah, Alex, et al.
Published: (2023)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
Exact Learning of Arithmetic with Differentiable Agents
by: Papazov, Hristo, et al.
Published: (2025)
by: Papazov, Hristo, et al.
Published: (2025)
Overparametrized linear dimensionality reductions: From projection pursuit to two-layer neural networks
by: Montanari, Andrea, et al.
Published: (2022)
by: Montanari, Andrea, et al.
Published: (2022)
The committee machine: Computational to statistical gaps in learning a two-layers neural network
by: Aubin, Benjamin, et al.
Published: (2018)
by: Aubin, Benjamin, et al.
Published: (2018)
Why Do We Need Weight Decay in Modern Deep Learning?
by: D'Angelo, Francesco, et al.
Published: (2023)
by: D'Angelo, Francesco, et al.
Published: (2023)
Optimization and generalization analysis for two-layer physics-informed neural networks without over-parametrization
by: Zeng, Zhihan, et al.
Published: (2025)
by: Zeng, Zhihan, et al.
Published: (2025)
Escaping mediocrity: how two-layer networks learn hard generalized linear models with SGD
by: Arnaboldi, Luca, et al.
Published: (2023)
by: Arnaboldi, Luca, et al.
Published: (2023)
Incremental Learning of Sparse Attention Patterns in Transformers
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
Long-Context Linear System Identification
by: Yüksel, Oğuz Kaan, et al.
Published: (2024)
by: Yüksel, Oğuz Kaan, et al.
Published: (2024)
Online Decision-Focused Learning
by: Capitaine, Aymeric, et al.
Published: (2025)
by: Capitaine, Aymeric, et al.
Published: (2025)
Asymptotics of feature learning in two-layer networks after one gradient-step
by: Cui, Hugo, et al.
Published: (2024)
by: Cui, Hugo, et al.
Published: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
by: Neuhaus, Yannic, et al.
Published: (2026)
by: Neuhaus, Yannic, et al.
Published: (2026)
Learning time-scales in two-layers neural networks
by: Berthier, Raphaël, et al.
Published: (2023)
by: Berthier, Raphaël, et al.
Published: (2023)
Similar Items
-
Simplicity bias and optimization threshold in two-layer ReLU networks
by: Boursier, Etienne, et al.
Published: (2024) -
Penalising the biases in norm regularisation enforces sparsity
by: Boursier, Etienne, et al.
Published: (2023) -
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
by: Boursier, Etienne, et al.
Published: (2022) -
First-order ANIL provably learns representations despite overparametrization
by: Yüksel, Oğuz Kaan, et al.
Published: (2023) -
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
by: Boursier, Etienne, et al.
Published: (2025)