Context-Gated Associative Retrieval: From Theory to Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Choraria, Moulik, Gerogiannis, Argyrios, Jayaraman, Vidhata, Mani, Ankur, Varshney, Lav R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
by: Hartman, Max, et al.
Published: (2026)
by: Hartman, Max, et al.
Published: (2026)
Exploring Loss Landscapes through the Lens of Spin Glass Theory
by: Liao, Hao, et al.
Published: (2024)
by: Liao, Hao, et al.
Published: (2024)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
Consciousness as a Jamming Phase
by: Ouyang, Kaichen
Published: (2025)
by: Ouyang, Kaichen
Published: (2025)
Analysis of Hopfield Model as Associative Memory
by: Silvestri, Matteo
Published: (2024)
by: Silvestri, Matteo
Published: (2024)
Applications of Statistical Field Theory in Deep Learning
by: Ringel, Zohar, et al.
Published: (2025)
by: Ringel, Zohar, et al.
Published: (2025)
From Embeddings to Dyson Series: Transformer Mechanics as Non-Hermitian Operator Theory
by: Chang, Po-Hao
Published: (2026)
by: Chang, Po-Hao
Published: (2026)
Parameter Symmetry Potentially Unifies Deep Learning Theory
by: Ziyin, Liu, et al.
Published: (2025)
by: Ziyin, Liu, et al.
Published: (2025)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
Approximation Theory for Neural Networks: Old and New
by: Mukherjee, Soumendu Sundar, et al.
Published: (2026)
by: Mukherjee, Soumendu Sundar, et al.
Published: (2026)
Do Hopfield Networks Dream of Stored Patterns? A Statistical-Mechanical Theory of Dreaming in Multidirectional Associative Memories
by: Barra, Adriano, et al.
Published: (2026)
by: Barra, Adriano, et al.
Published: (2026)
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
by: Lauditi, Clarissa, et al.
Published: (2026)
by: Lauditi, Clarissa, et al.
Published: (2026)
On the origin of neural scaling laws: from random graphs to natural language
by: Barkeshli, Maissam, et al.
Published: (2026)
by: Barkeshli, Maissam, et al.
Published: (2026)
Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model
by: Levi, Noam
Published: (2026)
by: Levi, Noam
Published: (2026)
More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)
by: Meir, Sagi, et al.
Published: (2026)
by: Meir, Sagi, et al.
Published: (2026)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
by: Defilippis, Leonardo, et al.
Published: (2025)
by: Defilippis, Leonardo, et al.
Published: (2025)
Generalization through variance: how noise shapes inductive biases in diffusion models
by: Vastola, John J.
Published: (2025)
by: Vastola, John J.
Published: (2025)
Identifying internal patterns in (1+1)-dimensional directed percolation using neural networks
by: Parkhomenko, Danil, et al.
Published: (2025)
by: Parkhomenko, Danil, et al.
Published: (2025)
Grokking vs. Learning: Same Features, Different Encodings
by: Manning-Coe, Dmitry, et al.
Published: (2025)
by: Manning-Coe, Dmitry, et al.
Published: (2025)
Alpha Zero for Physics: Application of Symbolic Regression with Alpha Zero to find the analytical methods in physics
by: Michishita, Yoshihiro
Published: (2023)
by: Michishita, Yoshihiro
Published: (2023)
A Spin Glass Characterization of Neural Networks
by: Li, Jun
Published: (2025)
by: Li, Jun
Published: (2025)
A method for quantifying the generalization capabilities of generative models for solving Ising models
by: Ma, Qunlong, et al.
Published: (2024)
by: Ma, Qunlong, et al.
Published: (2024)
KAN: Kolmogorov-Arnold Networks
by: Liu, Ziming, et al.
Published: (2024)
by: Liu, Ziming, et al.
Published: (2024)
Fast Analysis of the OpenAI O1-Preview Model in Solving Random K-SAT Problem: Does the LLM Solve the Problem Itself or Call an External SAT Solver?
by: Marino, Raffaele
Published: (2024)
by: Marino, Raffaele
Published: (2024)
Data-Driven Learnability Transition of Measurement-Induced Entanglement
by: Qian, Dongheng, et al.
Published: (2025)
by: Qian, Dongheng, et al.
Published: (2025)
A Geometric Perspective on the Difficulties of Learning GNN-based SAT Solvers
by: Skenderi, Geri
Published: (2025)
by: Skenderi, Geri
Published: (2025)
Dynamic Reservoir Computing with Physical Neuromorphic Networks
by: Xu, Yinhao, et al.
Published: (2025)
by: Xu, Yinhao, et al.
Published: (2025)
The Persian Rug: solving toy models of superposition using large-scale symmetries
by: Cowsik, Aditya, et al.
Published: (2024)
by: Cowsik, Aditya, et al.
Published: (2024)
Learning Chaotic Dynamics with Neuromorphic Network Dynamics
by: Xu, Yinhao, et al.
Published: (2025)
by: Xu, Yinhao, et al.
Published: (2025)
Enhancing Noise-Robust Losses for Large-Scale Noisy Data Learning
by: Staats, Max, et al.
Published: (2023)
by: Staats, Max, et al.
Published: (2023)
Representation Learning on a Random Lattice
by: Brill, Aryeh
Published: (2025)
by: Brill, Aryeh
Published: (2025)
Connecting NTK and NNGP: A Unified Theoretical Framework for Wide Neural Network Learning Dynamics
by: Avidan, Yehonatan, et al.
Published: (2023)
by: Avidan, Yehonatan, et al.
Published: (2023)
Towards Distributed Neural Architectures
by: Cowsik, Aditya, et al.
Published: (2025)
by: Cowsik, Aditya, et al.
Published: (2025)
Quantum-activated neural reservoirs on-chip open up large hardware security models for resilient authentication
by: He, Zhao, et al.
Published: (2024)
by: He, Zhao, et al.
Published: (2024)
Smooth Kolmogorov Arnold networks enabling structural knowledge representation
by: Samadi, Moein E., et al.
Published: (2024)
by: Samadi, Moein E., et al.
Published: (2024)
NCoder -- A Quantum Field Theory approach to encoding data
by: Berman, David S., et al.
Published: (2024)
by: Berman, David S., et al.
Published: (2024)
Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees
by: Breccia, Alessandro, et al.
Published: (2025)
by: Breccia, Alessandro, et al.
Published: (2025)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Similar Items
-
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
by: Hartman, Max, et al.
Published: (2026) -
Exploring Loss Landscapes through the Lens of Spin Glass Theory
by: Liao, Hao, et al.
Published: (2024) -
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025) -
Consciousness as a Jamming Phase
by: Ouyang, Kaichen
Published: (2025) -
Analysis of Hopfield Model as Associative Memory
by: Silvestri, Matteo
Published: (2024)