Self-attention as an attractor network: transient memories without backpropagation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | D'Amico, Francesco, Negri, Matteo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
von: D'Amico, Francesco, et al.
Veröffentlicht: (2025)
von: D'Amico, Francesco, et al.
Veröffentlicht: (2025)
Statistical mechanics of vector Hopfield network near and above saturation
von: Nicoletti, Flavio, et al.
Veröffentlicht: (2025)
von: Nicoletti, Flavio, et al.
Veröffentlicht: (2025)
Pseudo-likelihood produces associative memories able to generalize, even for asymmetric couplings
von: D'Amico, Francesco, et al.
Veröffentlicht: (2025)
von: D'Amico, Francesco, et al.
Veröffentlicht: (2025)
Benchmarking Graph Neural Networks in Solving Hard Constraint Satisfaction Problems
von: Skenderi, Geri, et al.
Veröffentlicht: (2026)
von: Skenderi, Geri, et al.
Veröffentlicht: (2026)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
von: Troiani, Emanuele, et al.
Veröffentlicht: (2025)
von: Troiani, Emanuele, et al.
Veröffentlicht: (2025)
Asymptotic theory of in-context learning by linear attention
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
Topological mechanical neural networks as classifiers through in situ backpropagation learning
von: Li, Shuaifeng, et al.
Veröffentlicht: (2025)
von: Li, Shuaifeng, et al.
Veröffentlicht: (2025)
Random Features Hopfield Networks generalize retrieval to previously unseen examples
von: Kalaj, Silvio, et al.
Veröffentlicht: (2024)
von: Kalaj, Silvio, et al.
Veröffentlicht: (2024)
How noise affects memory in linear recurrent networks
von: Guan, JingChuan, et al.
Veröffentlicht: (2024)
von: Guan, JingChuan, et al.
Veröffentlicht: (2024)
Specialization of softmax attention heads: insights from the high-dimensional single-location model
von: Sagitova, M., et al.
Veröffentlicht: (2026)
von: Sagitova, M., et al.
Veröffentlicht: (2026)
Non-equilibrium active noise enhances generative memory in diffusion models
von: Behera, Agnish Kumar, et al.
Veröffentlicht: (2024)
von: Behera, Agnish Kumar, et al.
Veröffentlicht: (2024)
Dynamical stability for dense patterns in discrete attractor neural networks
von: Cohen, Uri, et al.
Veröffentlicht: (2025)
von: Cohen, Uri, et al.
Veröffentlicht: (2025)
Exact full-RSB SAT/UNSAT transition in infinitely wide two-layer neural networks
von: Annesi, Brandon L., et al.
Veröffentlicht: (2024)
von: Annesi, Brandon L., et al.
Veröffentlicht: (2024)
Capacity of the Hebbian-Hopfield network associative memory
von: Stojnic, Mihailo
Veröffentlicht: (2024)
von: Stojnic, Mihailo
Veröffentlicht: (2024)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
High-dimensional learning of narrow neural networks
von: Cui, Hugo
Veröffentlicht: (2024)
von: Cui, Hugo
Veröffentlicht: (2024)
Emergent weight morphologies in deep neural networks
von: de Jong, Pascal, et al.
Veröffentlicht: (2025)
von: de Jong, Pascal, et al.
Veröffentlicht: (2025)
A universal approximation theorem for nonlinear resistive networks
von: Scellier, Benjamin, et al.
Veröffentlicht: (2023)
von: Scellier, Benjamin, et al.
Veröffentlicht: (2023)
Supervised and Unsupervised protocols for hetero-associative neural networks
von: Alessandrelli, Andrea, et al.
Veröffentlicht: (2025)
von: Alessandrelli, Andrea, et al.
Veröffentlicht: (2025)
Deep neural networks from the perspective of ergodic theory
von: Zhang, Fan
Veröffentlicht: (2023)
von: Zhang, Fan
Veröffentlicht: (2023)
Computing frustration and near-monotonicity in deep neural networks
von: Wendin, Joel, et al.
Veröffentlicht: (2025)
von: Wendin, Joel, et al.
Veröffentlicht: (2025)
Finite-time Lyapunov exponents of deep neural networks
von: Storm, L., et al.
Veröffentlicht: (2023)
von: Storm, L., et al.
Veröffentlicht: (2023)
A High Dimensional Statistical Model for Adversarial Training: Geometry and Trade-Offs
von: Tanner, Kasimir, et al.
Veröffentlicht: (2024)
von: Tanner, Kasimir, et al.
Veröffentlicht: (2024)
Homophily modulates double descent generalization in graph convolution networks
von: Shi, Cheng, et al.
Veröffentlicht: (2022)
von: Shi, Cheng, et al.
Veröffentlicht: (2022)
Training neural networks with structured noise improves classification and generalization
von: Benedetti, Marco, et al.
Veröffentlicht: (2023)
von: Benedetti, Marco, et al.
Veröffentlicht: (2023)
Modeling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting
von: Thériault, Robin, et al.
Veröffentlicht: (2024)
von: Thériault, Robin, et al.
Veröffentlicht: (2024)
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
von: Fioratti, Tommaso, et al.
Veröffentlicht: (2026)
von: Fioratti, Tommaso, et al.
Veröffentlicht: (2026)
Learning curves theory for hierarchically compositional data with power-law distributed features
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
Asymptotic generalization error of a single-layer graph convolutional network
von: Duranthon, O., et al.
Veröffentlicht: (2024)
von: Duranthon, O., et al.
Veröffentlicht: (2024)
How does training shape the Riemannian geometry of neural network representations?
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2023)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2023)
Asymptotics of feature learning in two-layer networks after one gradient-step
von: Cui, Hugo, et al.
Veröffentlicht: (2024)
von: Cui, Hugo, et al.
Veröffentlicht: (2024)
Boundary between noise and information applied to filtering neural network weight matrices
von: Staats, Max, et al.
Veröffentlicht: (2022)
von: Staats, Max, et al.
Veröffentlicht: (2022)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Bayes optimal learning of attention-indexed models
von: Boncoraglio, Fabrizio, et al.
Veröffentlicht: (2025)
von: Boncoraglio, Fabrizio, et al.
Veröffentlicht: (2025)
Neuronal correlations shape the scaling behavior of memory capacity and nonlinear computational capability of reservoir recurrent neural networks
von: Takasu, Shotaro, et al.
Veröffentlicht: (2025)
von: Takasu, Shotaro, et al.
Veröffentlicht: (2025)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
von: Dandi, Yatin, et al.
Veröffentlicht: (2026)
von: Dandi, Yatin, et al.
Veröffentlicht: (2026)
Towards a theory of how the structure of language is acquired by deep neural networks
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2024)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2024)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
von: Cagnetta, Francesco, et al.
Veröffentlicht: (2025)
Regularization, early-stopping and dreaming: a Hopfield-like setup to address generalization and overfitting
von: Agliari, Elena, et al.
Veröffentlicht: (2023)
von: Agliari, Elena, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
von: D'Amico, Francesco, et al.
Veröffentlicht: (2025) -
Statistical mechanics of vector Hopfield network near and above saturation
von: Nicoletti, Flavio, et al.
Veröffentlicht: (2025) -
Pseudo-likelihood produces associative memories able to generalize, even for asymmetric couplings
von: D'Amico, Francesco, et al.
Veröffentlicht: (2025) -
Benchmarking Graph Neural Networks in Solving Hard Constraint Satisfaction Problems
von: Skenderi, Geri, et al.
Veröffentlicht: (2026) -
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
von: Smart, Matthew, et al.
Veröffentlicht: (2025)