Attention layers provably solve single-location regression
Fuente:
arXiv
Saved in:
| Main Authors: | Marion, Pierre, Berthier, Raphaël, Biau, Gérard, Boyer, Claire |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
by: Wu, Yu-Han, et al.
Published: (2025)
by: Wu, Yu-Han, et al.
Published: (2025)
Optimal Stopping in Latent Diffusion Models
by: Wu, Yu-Han, et al.
Published: (2025)
by: Wu, Yu-Han, et al.
Published: (2025)
Physics-informed kernel learning
by: Doumèche, Nathan, et al.
Published: (2024)
by: Doumèche, Nathan, et al.
Published: (2024)
Fast kernel methods: Sobolev, physics-informed, and additive models
by: Doumèche, Nathan, et al.
Published: (2025)
by: Doumèche, Nathan, et al.
Published: (2025)
Attention-based clustering
by: Maulen-Soto, Rodrigo, et al.
Published: (2025)
by: Maulen-Soto, Rodrigo, et al.
Published: (2025)
Scaling ResNets in the Large-depth Regime
by: Marion, Pierre, et al.
Published: (2022)
by: Marion, Pierre, et al.
Published: (2022)
Implicit regularization of deep residual networks towards neural ODEs
by: Marion, Pierre, et al.
Published: (2023)
by: Marion, Pierre, et al.
Published: (2023)
A Note on k-NN Gating in RAG
by: Biau, Gérard, et al.
Published: (2026)
by: Biau, Gérard, et al.
Published: (2026)
Forecasting time series with constraints
by: Doumèche, Nathan, et al.
Published: (2025)
by: Doumèche, Nathan, et al.
Published: (2025)
Diagonal Linear Networks and the Lasso Regularization Path
by: Berthier, Raphaël
Published: (2025)
by: Berthier, Raphaël
Published: (2025)
Learning time-scales in two-layers neural networks
by: Berthier, Raphaël, et al.
Published: (2023)
by: Berthier, Raphaël, et al.
Published: (2023)
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
On the convergence of PINNs
by: Doumèche, Nathan, et al.
Published: (2023)
by: Doumèche, Nathan, et al.
Published: (2023)
A Geometry-Aware Residual Correction of Hagan's SABR Implied Volatility Formula
by: Reghai, Adil, et al.
Published: (2026)
by: Reghai, Adil, et al.
Published: (2026)
Attention-based PCA
by: Maulen-Soto, Rodrigo, et al.
Published: (2026)
by: Maulen-Soto, Rodrigo, et al.
Published: (2026)
Deep linear networks for regression are implicitly regularized towards flat minima
by: Marion, Pierre, et al.
Published: (2024)
by: Marion, Pierre, et al.
Published: (2024)
On the Minimal Degree Bias in Generalization on the Unseen for non-Boolean Functions
by: Pushkin, Denys, et al.
Published: (2024)
by: Pushkin, Denys, et al.
Published: (2024)
Towards graph neural networks for provably solving convex optimization problems
by: Qian, Chendi, et al.
Published: (2025)
by: Qian, Chendi, et al.
Published: (2025)
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
by: Zhu, Libin, et al.
Published: (2026)
by: Zhu, Libin, et al.
Published: (2026)
On provable privacy vulnerabilities of graph representations
by: Wu, Ruofan, et al.
Published: (2024)
by: Wu, Ruofan, et al.
Published: (2024)
Physics-informed machine learning as a kernel method
by: Doumèche, Nathan, et al.
Published: (2024)
by: Doumèche, Nathan, et al.
Published: (2024)
Des-q: a quantum algorithm to provably speedup retraining of decision trees
by: Kumar, Niraj, et al.
Published: (2023)
by: Kumar, Niraj, et al.
Published: (2023)
Multifidelity Gaussian process regression for solving nonlinear partial differential equations
by: El-Boukkouri, Fatima-Zahrae, et al.
Published: (2026)
by: El-Boukkouri, Fatima-Zahrae, et al.
Published: (2026)
A Hybrid Tsallis-Polarization Impurity Measure for Decision Trees: Theoretical Foundations and Empirical Evaluation
by: Lansiaux, Edouard, et al.
Published: (2026)
by: Lansiaux, Edouard, et al.
Published: (2026)
A randomized algorithm to solve reduced rank operator regression
by: Turri, Giacomo, et al.
Published: (2023)
by: Turri, Giacomo, et al.
Published: (2023)
Learning efficient and provably convergent splitting methods
by: Kreusser, L. M., et al.
Published: (2024)
by: Kreusser, L. M., et al.
Published: (2024)
Amortised and provably-robust simulation-based inference
by: Bharti, Ayush, et al.
Published: (2026)
by: Bharti, Ayush, et al.
Published: (2026)
First-order ANIL provably learns representations despite overparametrization
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
One-layer transformers fail to solve the induction heads task
by: Sanford, Clayton, et al.
Published: (2024)
by: Sanford, Clayton, et al.
Published: (2024)
Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
by: Duranthon, O., et al.
Published: (2025)
by: Duranthon, O., et al.
Published: (2025)
Linear Recursive Feature Machines provably recover low-rank matrices
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
Shallow diffusion networks provably learn hidden low-dimensional structure
by: Boffi, Nicholas M., et al.
Published: (2024)
by: Boffi, Nicholas M., et al.
Published: (2024)
Influence as soft sparsity: Estimation of monotone functions on $\{0,1\}^d$
by: Biau, Gérard
Published: (2026)
by: Biau, Gérard
Published: (2026)
Convergence Rates for Distribution Matching with Sliced Optimal Transport
by: Thurin, Gauthier, et al.
Published: (2026)
by: Thurin, Gauthier, et al.
Published: (2026)
Optimal Transport-based Conformal Prediction
by: Thurin, Gauthier, et al.
Published: (2025)
by: Thurin, Gauthier, et al.
Published: (2025)
Quantum automated learning with provable and explainable trainability
by: Ye, Qi, et al.
Published: (2025)
by: Ye, Qi, et al.
Published: (2025)
Does provable absence of barren plateaus imply classical simulability?
by: Cerezo, M., et al.
Published: (2023)
by: Cerezo, M., et al.
Published: (2023)
Dynamical simulation via quantum machine learning with provable generalization
by: Gibbs, Joe, et al.
Published: (2022)
by: Gibbs, Joe, et al.
Published: (2022)
Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents
by: Mangold, Paul, et al.
Published: (2025)
by: Mangold, Paul, et al.
Published: (2025)
Equivariant score-based generative models provably learn distributions with symmetries efficiently
by: Chen, Ziyu, et al.
Published: (2024)
by: Chen, Ziyu, et al.
Published: (2024)
Similar Items
-
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
by: Wu, Yu-Han, et al.
Published: (2025) -
Optimal Stopping in Latent Diffusion Models
by: Wu, Yu-Han, et al.
Published: (2025) -
Physics-informed kernel learning
by: Doumèche, Nathan, et al.
Published: (2024) -
Fast kernel methods: Sobolev, physics-informed, and additive models
by: Doumèche, Nathan, et al.
Published: (2025) -
Attention-based clustering
by: Maulen-Soto, Rodrigo, et al.
Published: (2025)