Fine-Tuning Dynamics of In-Context Factual Recall in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Ruomin, Nichani, Eshaan, Lee, Jason D., Ge, Rong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Factual Recall in Transformers via Associative Memories
by: Nichani, Eshaan, et al.
Published: (2024)
by: Nichani, Eshaan, et al.
Published: (2024)
Quantitative Bounds for Length Generalization in Transformers
by: Izzo, Zachary, et al.
Published: (2025)
by: Izzo, Zachary, et al.
Published: (2025)
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024)
by: Nichani, Eshaan, et al.
Published: (2024)
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
by: Nichani, Eshaan, et al.
Published: (2023)
by: Nichani, Eshaan, et al.
Published: (2023)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
Emergence and scaling laws in SGD learning of shallow neural networks
by: Ren, Yunwei, et al.
Published: (2025)
by: Ren, Yunwei, et al.
Published: (2025)
On the Statistical Query Complexity of Learning Semiautomata: a Random Walk Approach
by: Giapitzakis, George, et al.
Published: (2025)
by: Giapitzakis, George, et al.
Published: (2025)
Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks
by: Fu, Hengyu, et al.
Published: (2024)
by: Fu, Hengyu, et al.
Published: (2024)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
by: Kim, Juno, et al.
Published: (2026)
by: Kim, Juno, et al.
Published: (2026)
Learning Compositional Functions with Transformers from Easy-to-Hard Data
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval
by: Barnfield, Nicholas, et al.
Published: (2026)
by: Barnfield, Nicholas, et al.
Published: (2026)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
An Effective Dynamic Gradient Calibration Method for Continual Learning
by: Lin, Weichen, et al.
Published: (2024)
by: Lin, Weichen, et al.
Published: (2024)
How Transformers Learn In-Context Recall Tasks? Optimality, Training Dynamics and Generalization
by: Nguyen, Quan, et al.
Published: (2025)
by: Nguyen, Quan, et al.
Published: (2025)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
by: Chughtai, Bilal, et al.
Published: (2024)
by: Chughtai, Bilal, et al.
Published: (2024)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
by: Yuan, Jiaqing, et al.
Published: (2024)
by: Yuan, Jiaqing, et al.
Published: (2024)
Linear Transformers are Versatile In-Context Learners
by: Vladymyrov, Max, et al.
Published: (2024)
by: Vladymyrov, Max, et al.
Published: (2024)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
Decomposing Prediction Mechanisms for In-Context Recall
by: Daniels, Sultan, et al.
Published: (2025)
by: Daniels, Sultan, et al.
Published: (2025)
Locate-then-edit for Multi-hop Factual Recall under Knowledge Editing
by: Zhang, Zhuoran, et al.
Published: (2024)
by: Zhang, Zhuoran, et al.
Published: (2024)
ReCaLL: Membership Inference via Relative Conditional Log-Likelihoods
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
Dynamic Weight Grafting: Localizing Finetuned Factual Knowledge in Transformers
by: Nief, Todd, et al.
Published: (2025)
by: Nief, Todd, et al.
Published: (2025)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
by: Smith, Matthew L., et al.
Published: (2026)
by: Smith, Matthew L., et al.
Published: (2026)
Alignment Dynamics in LLM Fine-Tuning
by: Huang, Yuhan, et al.
Published: (2026)
by: Huang, Yuhan, et al.
Published: (2026)
Retrieval & Fine-Tuning for In-Context Tabular Models
by: Thomas, Valentin, et al.
Published: (2024)
by: Thomas, Valentin, et al.
Published: (2024)
LLM In-Context Recall is Prompt Dependent
by: Machlab, Daniel, et al.
Published: (2024)
by: Machlab, Daniel, et al.
Published: (2024)
Learning to Recall with Transformers Beyond Orthogonal Embeddings
by: Vural, Nuri Mert, et al.
Published: (2026)
by: Vural, Nuri Mert, et al.
Published: (2026)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
by: Lee, Chungpa, et al.
Published: (2026)
by: Lee, Chungpa, et al.
Published: (2026)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Fine-Tuned In-Context Learners for Efficient Adaptation
by: Bornschein, Jorg, et al.
Published: (2025)
by: Bornschein, Jorg, et al.
Published: (2025)
In-Context Learning and Fine-Tuning GPT for Argument Mining
by: Cabessa, Jérémie, et al.
Published: (2024)
by: Cabessa, Jérémie, et al.
Published: (2024)
Probably Approximately Precision and Recall Learning
by: Cohen, Lee, et al.
Published: (2024)
by: Cohen, Lee, et al.
Published: (2024)
Learning Orthogonal Multi-Index Models: A Fine-Grained Information Exponent Analysis
by: Ren, Yunwei, et al.
Published: (2024)
by: Ren, Yunwei, et al.
Published: (2024)
Optimizing DDPM Sampling with Shortcut Fine-Tuning
by: Fan, Ying, et al.
Published: (2023)
by: Fan, Ying, et al.
Published: (2023)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
by: Liu, James, et al.
Published: (2024)
by: Liu, James, et al.
Published: (2024)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
by: Lei, Ge, et al.
Published: (2025)
by: Lei, Ge, et al.
Published: (2025)
In-Context Probing for Membership Inference in Fine-Tuned Language Models
by: Lu, Zhexi, et al.
Published: (2025)
by: Lu, Zhexi, et al.
Published: (2025)
Tilt Matching for Scalable Sampling and Fine-Tuning
by: Potaptchik, Peter, et al.
Published: (2025)
by: Potaptchik, Peter, et al.
Published: (2025)
CF-VLM:CounterFactual Vision-Language Fine-tuning
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
by: Huang, Jianhao, et al.
Published: (2025)
by: Huang, Jianhao, et al.
Published: (2025)
Similar Items
-
Understanding Factual Recall in Transformers via Associative Memories
by: Nichani, Eshaan, et al.
Published: (2024) -
Quantitative Bounds for Length Generalization in Transformers
by: Izzo, Zachary, et al.
Published: (2025) -
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024) -
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
by: Nichani, Eshaan, et al.
Published: (2023) -
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)