Transformers are almost optimal metalearners for linear classification
Fuente:
arXiv
Saved in:
| Main Authors: | Magen, Roey, Vardi, Gal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Think from Multiple Thinkers
by: Joshi, Nirmit, et al.
Published: (2026)
by: Joshi, Nirmit, et al.
Published: (2026)
Benign Overfitting in Single-Head Attention
by: Magen, Roey, et al.
Published: (2024)
by: Magen, Roey, et al.
Published: (2024)
Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context
by: Frei, Spencer, et al.
Published: (2024)
by: Frei, Spencer, et al.
Published: (2024)
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
by: Gronich, Eitan, et al.
Published: (2026)
by: Gronich, Eitan, et al.
Published: (2026)
Are We Still Missing an Item?
by: Magen, Roey
Published: (2024)
by: Magen, Roey
Published: (2024)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
by: Joshi, Nirmit, et al.
Published: (2023)
by: Joshi, Nirmit, et al.
Published: (2023)
Implicit Regularization Towards Rank Minimization in ReLU Networks
by: Timor, Nadav, et al.
Published: (2022)
by: Timor, Nadav, et al.
Published: (2022)
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
by: Medvedev, Marko, et al.
Published: (2024)
by: Medvedev, Marko, et al.
Published: (2024)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
Provable Privacy Attacks on Trained Shallow Neural Networks
by: Smorodinsky, Guy, et al.
Published: (2024)
by: Smorodinsky, Guy, et al.
Published: (2024)
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
by: Tsilivis, Nikolaos, et al.
Published: (2024)
by: Tsilivis, Nikolaos, et al.
Published: (2024)
Provable Unlearning with Gradient Ascent on Two-Layer ReLU Neural Networks
by: Melamed, Odelia, et al.
Published: (2025)
by: Melamed, Odelia, et al.
Published: (2025)
An Agnostic View on the Cost of Overfitting in (Kernel) Ridge Regression
by: Zhou, Lijia, et al.
Published: (2023)
by: Zhou, Lijia, et al.
Published: (2023)
On the Hardness of Learning Regular Expressions
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
by: Harel, Itamar, et al.
Published: (2025)
by: Harel, Itamar, et al.
Published: (2025)
ImpMIA: Leveraging Implicit Bias for Membership Inference Attack
by: Golbari, Yuval, et al.
Published: (2025)
by: Golbari, Yuval, et al.
Published: (2025)
Positive Distribution Shift as a Framework for Understanding Tractable Learning
by: Medvedev, Marko, et al.
Published: (2026)
by: Medvedev, Marko, et al.
Published: (2026)
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
by: Bokobza, Roey, et al.
Published: (2024)
by: Bokobza, Roey, et al.
Published: (2024)
Provable Tempered Overfitting of Minimal Nets and Typical Nets
by: Harel, Itamar, et al.
Published: (2024)
by: Harel, Itamar, et al.
Published: (2024)
Flow matching achieves almost minimax optimal convergence
by: Fukumizu, Kenji, et al.
Published: (2024)
by: Fukumizu, Kenji, et al.
Published: (2024)
SOLVAR: Fast covariance-based heterogeneity analysis with pose refinement for cryo-EM
by: Yadgar, Roey, et al.
Published: (2026)
by: Yadgar, Roey, et al.
Published: (2026)
An efficient, provably optimal algorithm for the 0-1 loss linear classification problem
by: He, Xi, et al.
Published: (2023)
by: He, Xi, et al.
Published: (2023)
A Theory of Learning with Autoregressive Chain of Thought
by: Joshi, Nirmit, et al.
Published: (2025)
by: Joshi, Nirmit, et al.
Published: (2025)
Agnostic learning in (almost) optimal time via Gaussian surface area
by: Pesenti, Lucas, et al.
Published: (2026)
by: Pesenti, Lucas, et al.
Published: (2026)
Reconstructing Training Data From Real World Models Trained with Transfer Learning
by: Oz, Yakir, et al.
Published: (2024)
by: Oz, Yakir, et al.
Published: (2024)
Practical estimation of the optimal classification error with soft labels and calibration
by: Ushio, Ryota, et al.
Published: (2025)
by: Ushio, Ryota, et al.
Published: (2025)
On the optimal regret of collaborative personalized linear bandits
by: Huang, Bruce, et al.
Published: (2025)
by: Huang, Bruce, et al.
Published: (2025)
Pulse-based variational quantum optimization and metalearning in superconducting circuits
by: Wang, Yapeng, et al.
Published: (2024)
by: Wang, Yapeng, et al.
Published: (2024)
Sub-universal variational circuits for combinatorial optimization problems
by: Weitz, Gal, et al.
Published: (2023)
by: Weitz, Gal, et al.
Published: (2023)
Training Transformers for KV Cache Compressibility
by: Gelberg, Yoav, et al.
Published: (2026)
by: Gelberg, Yoav, et al.
Published: (2026)
Approaching Deep Learning through the Spectral Dynamics of Weights
by: Yunis, David, et al.
Published: (2024)
by: Yunis, David, et al.
Published: (2024)
Comparing Graph Transformers via Positional Encodings
by: Black, Mitchell, et al.
Published: (2024)
by: Black, Mitchell, et al.
Published: (2024)
PDE-SSM: A Spectral State Space Approach to Spatial Mixing in Diffusion Transformers
by: Gal, Eshed, et al.
Published: (2026)
by: Gal, Eshed, et al.
Published: (2026)
Inverse classification with logistic and softmax classifiers: efficient optimization
by: Carreira-Perpiñán, Miguel Á., et al.
Published: (2023)
by: Carreira-Perpiñán, Miguel Á., et al.
Published: (2023)
Sharp concentration of uniform generalization errors in binary linear classification
by: Nakakita, Shogo
Published: (2025)
by: Nakakita, Shogo
Published: (2025)
Fairness-aware Bayes optimal functional classification
by: Hu, Xiaoyu, et al.
Published: (2025)
by: Hu, Xiaoyu, et al.
Published: (2025)
Generating medical screening questionnaires through analysis of social media data
by: Ashkenazi, Ortal, et al.
Published: (2024)
by: Ashkenazi, Ortal, et al.
Published: (2024)
Adversarial Bandit over Bandits: Hierarchical Bandits for Online Configuration Management
by: Avin, Chen, et al.
Published: (2025)
by: Avin, Chen, et al.
Published: (2025)
Adversarial bandit optimization for approximately linear functions
by: Cheng, Zhuoyu, et al.
Published: (2025)
by: Cheng, Zhuoyu, et al.
Published: (2025)
Breaking the curse of dimensionality for linear rules: optimal predictors over the ellipsoid
by: Ayme, Alexis, et al.
Published: (2025)
by: Ayme, Alexis, et al.
Published: (2025)
Similar Items
-
Learning to Think from Multiple Thinkers
by: Joshi, Nirmit, et al.
Published: (2026) -
Benign Overfitting in Single-Head Attention
by: Magen, Roey, et al.
Published: (2024) -
Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context
by: Frei, Spencer, et al.
Published: (2024) -
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
by: Gronich, Eitan, et al.
Published: (2026) -
Are We Still Missing an Item?
by: Magen, Roey
Published: (2024)