On the rate of convergence of an over-parametrized Transformer classifier learned by gradient descent
Fuente:
arXiv
Saved in:
| Main Authors: | Kohler, Michael, Krzyzak, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analysis of the rate of convergence of an over-parametrized convolutional neural network image classifier learned by gradient descent
by: Kohler, Michael, et al.
Published: (2024)
by: Kohler, Michael, et al.
Published: (2024)
Learning of deep convolutional network image classifiers via stochastic gradient descent and over-parametrization
by: Kohler, Michael, et al.
Published: (2024)
by: Kohler, Michael, et al.
Published: (2024)
Statistically guided deep learning
by: Kohler, Michael, et al.
Published: (2025)
by: Kohler, Michael, et al.
Published: (2025)
On the rate of convergence of an over-parametrized deep neural network regression estimate learned by gradient descent
by: Kohler, Michael
Published: (2025)
by: Kohler, Michael
Published: (2025)
Spike-timing-dependent Hebbian learning as noisy gradient descent
by: Dexheimer, Niklas, et al.
Published: (2025)
by: Dexheimer, Niklas, et al.
Published: (2025)
On the rates of convergence for learning with convolutional neural networks
by: Yang, Yunfei, et al.
Published: (2024)
by: Yang, Yunfei, et al.
Published: (2024)
One-step corrected projected stochastic gradient descent for statistical estimation
by: Brouste, Alexandre, et al.
Published: (2023)
by: Brouste, Alexandre, et al.
Published: (2023)
Learning single index model with gradient descent: spectral initialization and precise asymptotics
by: Chen, Yuchen, et al.
Published: (2025)
by: Chen, Yuchen, et al.
Published: (2025)
Universality of high-dimensional scaling limits of stochastic gradient descent
by: Gheissari, Reza, et al.
Published: (2025)
by: Gheissari, Reza, et al.
Published: (2025)
Long-time dynamics and universality of nonconvex gradient descent
by: Han, Qiyang
Published: (2025)
by: Han, Qiyang
Published: (2025)
Fast Spawn\&Prune (FS\&P): Global convergence of stochastic conic particle gradient descent via birth/death process
by: De Castro, Yohann, et al.
Published: (2026)
by: De Castro, Yohann, et al.
Published: (2026)
On the existence of the maximum likelihood estimate and convergence rate under gradient descent for multi-class logistic regression
by: Nwaigwe, Dwight, et al.
Published: (2020)
by: Nwaigwe, Dwight, et al.
Published: (2020)
Convergence of flow-based generative models via proximal gradient descent in Wasserstein space
by: Cheng, Xiuyuan, et al.
Published: (2023)
by: Cheng, Xiuyuan, et al.
Published: (2023)
Uncertainty quantification by block bootstrap for differentially private stochastic gradient descent
by: Dette, Holger, et al.
Published: (2024)
by: Dette, Holger, et al.
Published: (2024)
Minimax rates of convergence for nonparametric regression under adversarial attacks
by: Peng, Jingfu, et al.
Published: (2024)
by: Peng, Jingfu, et al.
Published: (2024)
Online selective conformal inference: adaptive scores, convergence rate and optimality
by: Humbert, Pierre, et al.
Published: (2025)
by: Humbert, Pierre, et al.
Published: (2025)
Gradient descent for deep equilibrium single-index models
by: Dandapanthula, Sanjit, et al.
Published: (2025)
by: Dandapanthula, Sanjit, et al.
Published: (2025)
Building a stable classifier with the inflated argmax
by: Soloff, Jake A., et al.
Published: (2024)
by: Soloff, Jake A., et al.
Published: (2024)
Improved convergence rate of kNN graph Laplacians: differentiable self-tuned affinity
by: Cheng, Xiuyuan, et al.
Published: (2024)
by: Cheng, Xiuyuan, et al.
Published: (2024)
Contraction rates for conjugate gradient and Lanczos approximate posteriors in Gaussian process regression
by: Stankewitz, Bernhard, et al.
Published: (2024)
by: Stankewitz, Bernhard, et al.
Published: (2024)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Training thermodynamic computers by gradient descent
by: Whitelam, Stephen
Published: (2025)
by: Whitelam, Stephen
Published: (2025)
Implicit score matching meets denoising score matching: improved rates of convergence and log-density Hessian estimation
by: Yakovlev, Konstantin, et al.
Published: (2025)
by: Yakovlev, Konstantin, et al.
Published: (2025)
Sharp asymptotic theory for Q-learning with LDTZ learning rate and its generalization
by: Bonnerjee, Soham, et al.
Published: (2026)
by: Bonnerjee, Soham, et al.
Published: (2026)
Minimax optimal adaptive structured transfer learning through semi-parametric domain-varying coefficient model
by: Chen, Hanxiao, et al.
Published: (2026)
by: Chen, Hanxiao, et al.
Published: (2026)
Distributed Estimation and Inference for Semi-parametric Binary Response Models
by: Chen, Xi, et al.
Published: (2022)
by: Chen, Xi, et al.
Published: (2022)
Rates of convergence for density estimation with generative adversarial networks
by: Puchkin, Nikita, et al.
Published: (2021)
by: Puchkin, Nikita, et al.
Published: (2021)
eGAD! double descent is explained by Generalized Aliasing Decomposition
by: Transtrum, Mark K., et al.
Published: (2024)
by: Transtrum, Mark K., et al.
Published: (2024)
Neural Drift Estimation for Ergodic Diffusions: Non-parametric Analysis and Numerical Exploration
by: Di Gregorio, Simone, et al.
Published: (2025)
by: Di Gregorio, Simone, et al.
Published: (2025)
Adversarial learning for nonparametric regression: Minimax rate and adaptive estimation
by: Peng, Jingfu, et al.
Published: (2025)
by: Peng, Jingfu, et al.
Published: (2025)
Gradient descent inference in empirical risk minimization
by: Han, Qiyang, et al.
Published: (2024)
by: Han, Qiyang, et al.
Published: (2024)
Eigen-convergence of Gaussian kernelized graph Laplacian by manifold heat interpolation
by: Cheng, Xiuyuan, et al.
Published: (2021)
by: Cheng, Xiuyuan, et al.
Published: (2021)
Fast convergence of the Expectation Maximization algorithm under a logarithmic Sobolev inequality
by: Caprio, Rocco, et al.
Published: (2024)
by: Caprio, Rocco, et al.
Published: (2024)
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
by: Dragutinović, Sara, et al.
Published: (2025)
by: Dragutinović, Sara, et al.
Published: (2025)
Semi-parametric inference based on adaptively collected data
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
Accurate Evaluation of Quickest Changepoint Detectors via Non-parametric Survival Analysis
by: Miyagawa, Taiki, et al.
Published: (2026)
by: Miyagawa, Taiki, et al.
Published: (2026)
Sliced gradient-enhanced Kriging for high-dimensional function approximation
by: Cheng, Kai, et al.
Published: (2022)
by: Cheng, Kai, et al.
Published: (2022)
Bi-stochastically normalized graph Laplacian: convergence to manifold Laplacian and robustness to outlier noise
by: Cheng, Xiuyuan, et al.
Published: (2022)
by: Cheng, Xiuyuan, et al.
Published: (2022)
Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer
by: Mustafi, Aratrika, et al.
Published: (2026)
by: Mustafi, Aratrika, et al.
Published: (2026)
Asymptotic spectrum of weighted sample covariance: another proof of spectrum convergence
by: Oriol, Benoit
Published: (2024)
by: Oriol, Benoit
Published: (2024)
Similar Items
-
Analysis of the rate of convergence of an over-parametrized convolutional neural network image classifier learned by gradient descent
by: Kohler, Michael, et al.
Published: (2024) -
Learning of deep convolutional network image classifiers via stochastic gradient descent and over-parametrization
by: Kohler, Michael, et al.
Published: (2024) -
Statistically guided deep learning
by: Kohler, Michael, et al.
Published: (2025) -
On the rate of convergence of an over-parametrized deep neural network regression estimate learned by gradient descent
by: Kohler, Michael
Published: (2025) -
Spike-timing-dependent Hebbian learning as noisy gradient descent
by: Dexheimer, Niklas, et al.
Published: (2025)