First-order ANIL provably learns representations despite overparametrization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yüksel, Oğuz Kaan, Boursier, Etienne, Flammarion, Nicolas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Penalising the biases in norm regularisation enforces sparsity
von: Boursier, Etienne, et al.
Veröffentlicht: (2023)
von: Boursier, Etienne, et al.
Veröffentlicht: (2023)
Early alignment in two-layer networks training is a two-edged sword
von: Boursier, Etienne, et al.
Veröffentlicht: (2024)
von: Boursier, Etienne, et al.
Veröffentlicht: (2024)
Simplicity bias and optimization threshold in two-layer ReLU networks
von: Boursier, Etienne, et al.
Veröffentlicht: (2024)
von: Boursier, Etienne, et al.
Veröffentlicht: (2024)
Long-Context Linear System Identification
von: Yüksel, Oğuz Kaan, et al.
Veröffentlicht: (2024)
von: Yüksel, Oğuz Kaan, et al.
Veröffentlicht: (2024)
Incremental Learning of Sparse Attention Patterns in Transformers
von: Yüksel, Oğuz Kaan, et al.
Veröffentlicht: (2026)
von: Yüksel, Oğuz Kaan, et al.
Veröffentlicht: (2026)
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
von: Boursier, Etienne, et al.
Veröffentlicht: (2022)
von: Boursier, Etienne, et al.
Veröffentlicht: (2022)
Benignity of loss landscape with weight decay requires both large overparametrization and initialization
von: Boursier, Etienne, et al.
Veröffentlicht: (2025)
von: Boursier, Etienne, et al.
Veröffentlicht: (2025)
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
von: Boursier, Etienne, et al.
Veröffentlicht: (2025)
von: Boursier, Etienne, et al.
Veröffentlicht: (2025)
On provable privacy vulnerabilities of graph representations
von: Wu, Ruofan, et al.
Veröffentlicht: (2024)
von: Wu, Ruofan, et al.
Veröffentlicht: (2024)
A survey on multi-player bandits
von: Boursier, Etienne, et al.
Veröffentlicht: (2022)
von: Boursier, Etienne, et al.
Veröffentlicht: (2022)
Transformers Learn Latent Mixture Models In-Context via Mirror Descent
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2026)
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2026)
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
von: Boursier, Etienne, et al.
Veröffentlicht: (2025)
von: Boursier, Etienne, et al.
Veröffentlicht: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points
von: Varre, Aditya, et al.
Veröffentlicht: (2025)
von: Varre, Aditya, et al.
Veröffentlicht: (2025)
Learning Parametric Distributions from Samples and Preferences
von: Jourdan, Marc, et al.
Veröffentlicht: (2025)
von: Jourdan, Marc, et al.
Veröffentlicht: (2025)
Effects of noise on the overparametrization of quantum neural networks
von: García-Martín, Diego, et al.
Veröffentlicht: (2023)
von: García-Martín, Diego, et al.
Veröffentlicht: (2023)
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
von: Papazov, Hristo, et al.
Veröffentlicht: (2024)
von: Papazov, Hristo, et al.
Veröffentlicht: (2024)
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
von: Varre, Aditya, et al.
Veröffentlicht: (2026)
von: Varre, Aditya, et al.
Veröffentlicht: (2026)
(How) Learning Rates Regulate Catastrophic Overtraining
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias
von: Town, James, et al.
Veröffentlicht: (2026)
von: Town, James, et al.
Veröffentlicht: (2026)
Tractability from overparametrization: The example of the negative perceptron
von: Montanari, Andrea, et al.
Veröffentlicht: (2021)
von: Montanari, Andrea, et al.
Veröffentlicht: (2021)
Quantum automated learning with provable and explainable trainability
von: Ye, Qi, et al.
Veröffentlicht: (2025)
von: Ye, Qi, et al.
Veröffentlicht: (2025)
Shallow diffusion networks provably learn hidden low-dimensional structure
von: Boffi, Nicholas M., et al.
Veröffentlicht: (2024)
von: Boffi, Nicholas M., et al.
Veröffentlicht: (2024)
Selective Induction Heads: How Transformers Select Causal Structures In Context
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2025)
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2025)
Learning Algorithms in the Limit
von: Papazov, Hristo, et al.
Veröffentlicht: (2025)
von: Papazov, Hristo, et al.
Veröffentlicht: (2025)
Dynamical simulation via quantum machine learning with provable generalization
von: Gibbs, Joe, et al.
Veröffentlicht: (2022)
von: Gibbs, Joe, et al.
Veröffentlicht: (2022)
Approximate information maximization for bandit games
von: Barbier-Chebbah, Alex, et al.
Veröffentlicht: (2023)
von: Barbier-Chebbah, Alex, et al.
Veröffentlicht: (2023)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Exact Learning of Arithmetic with Differentiable Agents
von: Papazov, Hristo, et al.
Veröffentlicht: (2025)
von: Papazov, Hristo, et al.
Veröffentlicht: (2025)
Implicit Bias of Mirror Flow on Separable Data
von: Pesme, Scott, et al.
Veröffentlicht: (2024)
von: Pesme, Scott, et al.
Veröffentlicht: (2024)
ODE approximation for the Adam algorithm: General and overparametrized setting
von: Dereich, Steffen, et al.
Veröffentlicht: (2025)
von: Dereich, Steffen, et al.
Veröffentlicht: (2025)
Entanglement-induced provable and robust quantum learning advantages
von: Zhao, Haimeng, et al.
Veröffentlicht: (2024)
von: Zhao, Haimeng, et al.
Veröffentlicht: (2024)
Unrolled denoising networks provably learn optimal Bayesian inference
von: Karan, Aayush, et al.
Veröffentlicht: (2024)
von: Karan, Aayush, et al.
Veröffentlicht: (2024)
Normalizing self-supervised learning for provably reliable Change Point Detection
von: Bazarova, Alexandra, et al.
Veröffentlicht: (2024)
von: Bazarova, Alexandra, et al.
Veröffentlicht: (2024)
Why Do We Need Weight Decay in Modern Deep Learning?
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
Equivariant score-based generative models provably learn distributions with symmetries efficiently
von: Chen, Ziyu, et al.
Veröffentlicht: (2024)
von: Chen, Ziyu, et al.
Veröffentlicht: (2024)
Towards provably efficient quantum algorithms for large-scale machine-learning models
von: Liu, Junyu, et al.
Veröffentlicht: (2023)
von: Liu, Junyu, et al.
Veröffentlicht: (2023)
Online Decision-Focused Learning
von: Capitaine, Aymeric, et al.
Veröffentlicht: (2025)
von: Capitaine, Aymeric, et al.
Veröffentlicht: (2025)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
von: Schlarmann, Christian, et al.
Veröffentlicht: (2025)
von: Schlarmann, Christian, et al.
Veröffentlicht: (2025)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Penalising the biases in norm regularisation enforces sparsity
von: Boursier, Etienne, et al.
Veröffentlicht: (2023) -
Early alignment in two-layer networks training is a two-edged sword
von: Boursier, Etienne, et al.
Veröffentlicht: (2024) -
Simplicity bias and optimization threshold in two-layer ReLU networks
von: Boursier, Etienne, et al.
Veröffentlicht: (2024) -
Long-Context Linear System Identification
von: Yüksel, Oğuz Kaan, et al.
Veröffentlicht: (2024) -
Incremental Learning of Sparse Attention Patterns in Transformers
von: Yüksel, Oğuz Kaan, et al.
Veröffentlicht: (2026)