In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Smart, Matthew, Bietti, Alberto, Sengupta, Anirvan M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asymptotic theory of in-context learning by linear attention
by: Lu, Yue M., et al.
Published: (2024)
by: Lu, Yue M., et al.
Published: (2024)
Self-attention as an attractor network: transient memories without backpropagation
by: D'Amico, Francesco, et al.
Published: (2024)
by: D'Amico, Francesco, et al.
Published: (2024)
Asymptotics of feature learning in two-layer networks after one gradient-step
by: Cui, Hugo, et al.
Published: (2024)
by: Cui, Hugo, et al.
Published: (2024)
Finite-size scaling of hetero-associative retrieval in continuous-signal-driven Ising spin systems
by: Ladiana, Andrea
Published: (2026)
by: Ladiana, Andrea
Published: (2026)
Solution space and storage capacity of fully connected two-layer neural networks with generic activation functions
by: Nishiyama, Sota, et al.
Published: (2024)
by: Nishiyama, Sota, et al.
Published: (2024)
Distinct mechanisms underlying in-context learning in transformers
by: Gibson, Cole, et al.
Published: (2026)
by: Gibson, Cole, et al.
Published: (2026)
Specialization of softmax attention heads: insights from the high-dimensional single-location model
by: Sagitova, M., et al.
Published: (2026)
by: Sagitova, M., et al.
Published: (2026)
Random Features Hopfield Networks generalize retrieval to previously unseen examples
by: Kalaj, Silvio, et al.
Published: (2024)
by: Kalaj, Silvio, et al.
Published: (2024)
Topological Exploration of High-Dimensional Empirical Risk Landscapes: general approach, and applications to phase retrieval
by: Maillard, Antoine, et al.
Published: (2026)
by: Maillard, Antoine, et al.
Published: (2026)
Thermodynamics of bidirectional associative memories
by: Barra, Adriano, et al.
Published: (2022)
by: Barra, Adriano, et al.
Published: (2022)
Non-equilibrium active noise enhances generative memory in diffusion models
by: Behera, Agnish Kumar, et al.
Published: (2024)
by: Behera, Agnish Kumar, et al.
Published: (2024)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
by: Troiani, Emanuele, et al.
Published: (2025)
by: Troiani, Emanuele, et al.
Published: (2025)
Properties of the geometry of solutions and capacity of multi-layer neural networks with Rectified Linear Units activations
by: Baldassi, Carlo, et al.
Published: (2019)
by: Baldassi, Carlo, et al.
Published: (2019)
Factual recall in linear associative memories: sharp asymptotics and mechanistic insights
by: Giorlandino, Alessio, et al.
Published: (2026)
by: Giorlandino, Alessio, et al.
Published: (2026)
Asymptotic generalization error of a single-layer graph convolutional network
by: Duranthon, O., et al.
Published: (2024)
by: Duranthon, O., et al.
Published: (2024)
On the phase diagram of extensive-rank symmetric matrix denoising beyond rotational invariance
by: Barbier, Jean, et al.
Published: (2024)
by: Barbier, Jean, et al.
Published: (2024)
Pseudo-likelihood produces associative memories able to generalize, even for asymmetric couplings
by: D'Amico, Francesco, et al.
Published: (2025)
by: D'Amico, Francesco, et al.
Published: (2025)
Perfect reconstruction of sparse signals using nonconvexity control and one-step RSB message passing
by: Gu, Xiaosi, et al.
Published: (2025)
by: Gu, Xiaosi, et al.
Published: (2025)
Capacity of the Hebbian-Hopfield network associative memory
by: Stojnic, Mihailo
Published: (2024)
by: Stojnic, Mihailo
Published: (2024)
Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks
by: Huang, Jie, et al.
Published: (2026)
by: Huang, Jie, et al.
Published: (2026)
Supervised and Unsupervised protocols for hetero-associative neural networks
by: Alessandrelli, Andrea, et al.
Published: (2025)
by: Alessandrelli, Andrea, et al.
Published: (2025)
Generalization performance of narrow one-hidden layer networks in the teacher-student setting
by: Ortiz, Rodrigo Pérez, et al.
Published: (2025)
by: Ortiz, Rodrigo Pérez, et al.
Published: (2025)
A Federated Many-to-One Hopfield model for associative Neural Networks
by: Alessandrelli, Andrea, et al.
Published: (2026)
by: Alessandrelli, Andrea, et al.
Published: (2026)
Pruning-induced phases in fully-connected neural networks: the eumentia, the dementia, and the amentia
by: Pan, Haining, et al.
Published: (2026)
by: Pan, Haining, et al.
Published: (2026)
How noise affects memory in linear recurrent networks
by: Guan, JingChuan, et al.
Published: (2024)
by: Guan, JingChuan, et al.
Published: (2024)
Bayes optimal learning of attention-indexed models
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
Exact full-RSB SAT/UNSAT transition in infinitely wide two-layer neural networks
by: Annesi, Brandon L., et al.
Published: (2024)
by: Annesi, Brandon L., et al.
Published: (2024)
Boundary between noise and information applied to filtering neural network weight matrices
by: Staats, Max, et al.
Published: (2022)
by: Staats, Max, et al.
Published: (2022)
Regularization, early-stopping and dreaming: a Hopfield-like setup to address generalization and overfitting
by: Agliari, Elena, et al.
Published: (2023)
by: Agliari, Elena, et al.
Published: (2023)
Superconductivity in the two-dimensional Hubbard model revealed by neural quantum states
by: Roth, Christopher, et al.
Published: (2025)
by: Roth, Christopher, et al.
Published: (2025)
Comparing the effects of Boltzmann machines as associative memory in Generative Adversarial Networks between classical and quantum sampling
by: Urushibata, Mitsuru, et al.
Published: (2022)
by: Urushibata, Mitsuru, et al.
Published: (2022)
A solvable model of learning generative diffusion: theory and insights
by: Cui, Hugo, et al.
Published: (2025)
by: Cui, Hugo, et al.
Published: (2025)
Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model
by: Pezzicoli, F. S., et al.
Published: (2025)
by: Pezzicoli, F. S., et al.
Published: (2025)
Deep networks learn to parse uniform-depth context-free languages from local statistics
by: Parley, Jack T., et al.
Published: (2026)
by: Parley, Jack T., et al.
Published: (2026)
Physics-inspired transformer quantum states via latent imaginary-time evolution
by: Yamazaki, Kimihiro, et al.
Published: (2026)
by: Yamazaki, Kimihiro, et al.
Published: (2026)
Neuronal correlations shape the scaling behavior of memory capacity and nonlinear computational capability of reservoir recurrent neural networks
by: Takasu, Shotaro, et al.
Published: (2025)
by: Takasu, Shotaro, et al.
Published: (2025)
Investigating layer-selective transfer learning of QAOA parameters for Max-Cut problem
by: Venturelli, Francesco Aldo, et al.
Published: (2024)
by: Venturelli, Francesco Aldo, et al.
Published: (2024)
Escape dynamics and implicit bias of one-pass SGD in overparameterized quadratic networks
by: Bocchi, Dario, et al.
Published: (2026)
by: Bocchi, Dario, et al.
Published: (2026)
Transparency versus Anderson localization in one-dimensional disordered stealthy hyperuniform layered media
by: Klatt, Michael A., et al.
Published: (2025)
by: Klatt, Michael A., et al.
Published: (2025)
Double descent: When do neural quantum states generalize?
by: Moss, M. Schuyler, et al.
Published: (2025)
by: Moss, M. Schuyler, et al.
Published: (2025)
Similar Items
-
Asymptotic theory of in-context learning by linear attention
by: Lu, Yue M., et al.
Published: (2024) -
Self-attention as an attractor network: transient memories without backpropagation
by: D'Amico, Francesco, et al.
Published: (2024) -
Asymptotics of feature learning in two-layer networks after one gradient-step
by: Cui, Hugo, et al.
Published: (2024) -
Finite-size scaling of hetero-associative retrieval in continuous-signal-driven Ising spin systems
by: Ladiana, Andrea
Published: (2026) -
Solution space and storage capacity of fully connected two-layer neural networks with generic activation functions
by: Nishiyama, Sota, et al.
Published: (2024)