Understanding Factual Recall in Transformers via Associative Memories
Fuente:
arXiv
Saved in:
| Main Authors: | Nichani, Eshaan, Lee, Jason D., Bietti, Alberto |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024)
by: Nichani, Eshaan, et al.
Published: (2024)
Fine-Tuning Dynamics of In-Context Factual Recall in Transformers
by: Huang, Ruomin, et al.
Published: (2026)
by: Huang, Ruomin, et al.
Published: (2026)
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval
by: Barnfield, Nicholas, et al.
Published: (2026)
by: Barnfield, Nicholas, et al.
Published: (2026)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
by: Kim, Juno, et al.
Published: (2026)
by: Kim, Juno, et al.
Published: (2026)
Geometric Factual Recall in Transformers
by: Ravfogel, Shauli, et al.
Published: (2026)
by: Ravfogel, Shauli, et al.
Published: (2026)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Quantitative Bounds for Length Generalization in Transformers
by: Izzo, Zachary, et al.
Published: (2025)
by: Izzo, Zachary, et al.
Published: (2025)
Learning Compositional Functions with Transformers from Easy-to-Hard Data
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks
by: Fu, Hengyu, et al.
Published: (2024)
by: Fu, Hengyu, et al.
Published: (2024)
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
by: Nichani, Eshaan, et al.
Published: (2023)
by: Nichani, Eshaan, et al.
Published: (2023)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
Quantifying Logical Consistency in Transformers via Query-Key Alignment
by: Tulchinskii, Eduard, et al.
Published: (2025)
by: Tulchinskii, Eduard, et al.
Published: (2025)
From Topic to Transition Structure: Unsupervised Concept Discovery at Corpus Scale via Predictive Associative Memory
by: Dury, Jason
Published: (2026)
by: Dury, Jason
Published: (2026)
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
by: Fesharaki, Amirmehdi Jafari, et al.
Published: (2026)
by: Fesharaki, Amirmehdi Jafari, et al.
Published: (2026)
An Information-theoretic Multi-task Representation Learning Framework for Natural Language Understanding
by: Hu, Dou, et al.
Published: (2025)
by: Hu, Dou, et al.
Published: (2025)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
An Enhanced Text Compression Approach Using Transformer-based Language Models
by: Rahman, Chowdhury Mofizur, et al.
Published: (2024)
by: Rahman, Chowdhury Mofizur, et al.
Published: (2024)
Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
by: Latimer, Chris, et al.
Published: (2025)
by: Latimer, Chris, et al.
Published: (2025)
Scaling Laws for Associative Memories
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Emergence and scaling laws in SGD learning of shallow neural networks
by: Ren, Yunwei, et al.
Published: (2025)
by: Ren, Yunwei, et al.
Published: (2025)
On the Statistical Query Complexity of Learning Semiautomata: a Random Walk Approach
by: Giapitzakis, George, et al.
Published: (2025)
by: Giapitzakis, George, et al.
Published: (2025)
Towards Optimal Statistical Watermarking
by: Huang, Baihe, et al.
Published: (2023)
by: Huang, Baihe, et al.
Published: (2023)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
by: Chughtai, Bilal, et al.
Published: (2024)
by: Chughtai, Bilal, et al.
Published: (2024)
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
by: Ma, Huidong, et al.
Published: (2026)
by: Ma, Huidong, et al.
Published: (2026)
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
by: Zuo, Fei, et al.
Published: (2026)
by: Zuo, Fei, et al.
Published: (2026)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
by: Ardestani, MohamamdJavad, et al.
Published: (2025)
by: Ardestani, MohamamdJavad, et al.
Published: (2025)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
Microstructures and Accuracy of Graph Recall by Large Language Models
by: Wang, Yanbang, et al.
Published: (2024)
by: Wang, Yanbang, et al.
Published: (2024)
InfAlign: Inference-aware language model alignment
by: Balashankar, Ananth, et al.
Published: (2024)
by: Balashankar, Ananth, et al.
Published: (2024)
Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Theoretical guarantees on the best-of-n alignment policy
by: Beirami, Ahmad, et al.
Published: (2024)
by: Beirami, Ahmad, et al.
Published: (2024)
Proposal and study of statistical features for string similarity computation and classification
by: Rodrigues, E. O., et al.
Published: (2026)
by: Rodrigues, E. O., et al.
Published: (2026)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
by: Krasnovsky, Anatoly A.
Published: (2025)
by: Krasnovsky, Anatoly A.
Published: (2025)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
by: Yuan, Jiaqing, et al.
Published: (2024)
by: Yuan, Jiaqing, et al.
Published: (2024)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
by: Wieser, Frederico, et al.
Published: (2025)
by: Wieser, Frederico, et al.
Published: (2025)
Integrating Pre-Trained Language Model with Physical Layer Communications
by: Lee, Ju-Hyung, et al.
Published: (2024)
by: Lee, Ju-Hyung, et al.
Published: (2024)
Does Privacy Always Harm Fairness? Data-Dependent Trade-offs via Chernoff Information Neural Estimation
by: Nichani, Arjun, et al.
Published: (2026)
by: Nichani, Arjun, et al.
Published: (2026)
Similar Items
-
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024) -
Fine-Tuning Dynamics of In-Context Factual Recall in Transformers
by: Huang, Ruomin, et al.
Published: (2026) -
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval
by: Barnfield, Nicholas, et al.
Published: (2026) -
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
by: Kim, Juno, et al.
Published: (2026) -
Geometric Factual Recall in Transformers
by: Ravfogel, Shauli, et al.
Published: (2026)