Asymptotic theory of in-context learning by linear attention
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Yue M., Letey, Mary I., Zavatone-Veth, Jacob A., Maiti, Anindita, Pehlevan, Cengiz |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Nadaraya-Watson kernel smoothing as a random energy model
by: Zavatone-Veth, Jacob A., et al.
Published: (2024)
by: Zavatone-Veth, Jacob A., et al.
Published: (2024)
A note on the dynamics of extended-context disordered kinetic spin models
by: Zavatone-Veth, Jacob A., et al.
Published: (2025)
by: Zavatone-Veth, Jacob A., et al.
Published: (2025)
Risk and cross validation in ridge regression with correlated samples
by: Atanasov, Alexander, et al.
Published: (2024)
by: Atanasov, Alexander, et al.
Published: (2024)
Scaling and renormalization in high-dimensional regression
by: Atanasov, Alexander, et al.
Published: (2024)
by: Atanasov, Alexander, et al.
Published: (2024)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
How does training shape the Riemannian geometry of neural network representations?
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
Dynamically Learning to Integrate in Recurrent Neural Networks
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
A solvable model of learning generative diffusion: theory and insights
by: Cui, Hugo, et al.
Published: (2025)
by: Cui, Hugo, et al.
Published: (2025)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
How Feature Learning Can Improve Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Transfer Learning in Infinite Width Feature Learning Networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
A Dynamical Model of Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Infinite Limits of Multi-head Transformer Dynamics
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
by: Ruben, Benjamin S., et al.
Published: (2023)
by: Ruben, Benjamin S., et al.
Published: (2023)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
by: Ruben, Benjamin S., et al.
Published: (2024)
by: Ruben, Benjamin S., et al.
Published: (2024)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
by: Lauditi, Clarissa, et al.
Published: (2026)
by: Lauditi, Clarissa, et al.
Published: (2026)
Exploring the Energy Landscape of RBMs: Reciprocal Space Insights into Bosons, Hierarchical Learning and Symmetry Breaking
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
Asymptotics of feature learning in two-layer networks after one gradient-step
by: Cui, Hugo, et al.
Published: (2024)
by: Cui, Hugo, et al.
Published: (2024)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
by: Smart, Matthew, et al.
Published: (2025)
by: Smart, Matthew, et al.
Published: (2025)
Applying statistical learning theory to deep learning
by: Gerbelot, Cédric, et al.
Published: (2023)
by: Gerbelot, Cédric, et al.
Published: (2023)
Pretrain-Test Task Alignment Governs Generalization in In-Context Learning
by: Letey, Mary I., et al.
Published: (2025)
by: Letey, Mary I., et al.
Published: (2025)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
by: Troiani, Emanuele, et al.
Published: (2025)
by: Troiani, Emanuele, et al.
Published: (2025)
A unified theory of feature learning in RNNs and DNNs
by: Bauer, Jan P., et al.
Published: (2026)
by: Bauer, Jan P., et al.
Published: (2026)
Specialization of softmax attention heads: insights from the high-dimensional single-location model
by: Sagitova, M., et al.
Published: (2026)
by: Sagitova, M., et al.
Published: (2026)
Wilsonian Renormalization of Neural Network Gaussian Processes
by: Howard, Jessica N., et al.
Published: (2024)
by: Howard, Jessica N., et al.
Published: (2024)
Bayes optimal learning of attention-indexed models
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
Self-attention as an attractor network: transient memories without backpropagation
by: D'Amico, Francesco, et al.
Published: (2024)
by: D'Amico, Francesco, et al.
Published: (2024)
Bayesian RG Flow in Neural Network Field Theories
by: Howard, Jessica N., et al.
Published: (2024)
by: Howard, Jessica N., et al.
Published: (2024)
High-dimensional Asymptotics of Denoising Autoencoders
by: Cui, Hugo, et al.
Published: (2023)
by: Cui, Hugo, et al.
Published: (2023)
Distinct mechanisms underlying in-context learning in transformers
by: Gibson, Cole, et al.
Published: (2026)
by: Gibson, Cole, et al.
Published: (2026)
Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
by: Xu, Yizhou, et al.
Published: (2025)
by: Xu, Yizhou, et al.
Published: (2025)
Asymptotic generalization error of a single-layer graph convolutional network
by: Duranthon, O., et al.
Published: (2024)
by: Duranthon, O., et al.
Published: (2024)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
by: Arnaboldi, Luca, et al.
Published: (2025)
by: Arnaboldi, Luca, et al.
Published: (2025)
Deep neural networks from the perspective of ergodic theory
by: Zhang, Fan
Published: (2023)
by: Zhang, Fan
Published: (2023)
Field theory for optimal signal propagation in ResNets
by: Fischer, Kirsten, et al.
Published: (2023)
by: Fischer, Kirsten, et al.
Published: (2023)
High-dimensional learning of narrow neural networks
by: Cui, Hugo
Published: (2024)
by: Cui, Hugo
Published: (2024)
Similar Items
-
Nadaraya-Watson kernel smoothing as a random energy model
by: Zavatone-Veth, Jacob A., et al.
Published: (2024) -
A note on the dynamics of extended-context disordered kinetic spin models
by: Zavatone-Veth, Jacob A., et al.
Published: (2025) -
Risk and cross validation in ridge regression with correlated samples
by: Atanasov, Alexander, et al.
Published: (2024) -
Scaling and renormalization in high-dimensional regression
by: Atanasov, Alexander, et al.
Published: (2024) -
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
by: Bordelon, Blake, et al.
Published: (2025)