Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Ying, Heinzerling, Benjamin, Li, Dongyuan, Inui, Kentaro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Monotonic Representation of Numeric Properties in Language Models
von: Heinzerling, Benjamin, et al.
Veröffentlicht: (2024)
von: Heinzerling, Benjamin, et al.
Veröffentlicht: (2024)
Representational Analysis of Binding in Language Models
von: Dai, Qin, et al.
Veröffentlicht: (2024)
von: Dai, Qin, et al.
Veröffentlicht: (2024)
Cell-Based Representation of Relational Binding in Language Models
von: Dai, Qin, et al.
Veröffentlicht: (2026)
von: Dai, Qin, et al.
Veröffentlicht: (2026)
Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference
von: Kamoda, Go, et al.
Veröffentlicht: (2025)
von: Kamoda, Go, et al.
Veröffentlicht: (2025)
TopK Language Models
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2025)
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2025)
The Curse of Popularity: Popular Entities have Catastrophic Side Effects when Deleting Knowledge from Language Models
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2024)
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2024)
On Entity Identification in Language Models
von: Sakata, Masaki, et al.
Veröffentlicht: (2025)
von: Sakata, Masaki, et al.
Veröffentlicht: (2025)
Linear Representations of Hierarchical Concepts in Language Models
von: Sakata, Masaki, et al.
Veröffentlicht: (2026)
von: Sakata, Masaki, et al.
Veröffentlicht: (2026)
What Matters in Memorizing and Recalling Facts? Multifaceted Benchmarks for Knowledge Probing in Language Models
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
von: El-Shangiti, Ahmed Oumar, et al.
Veröffentlicht: (2024)
von: El-Shangiti, Ahmed Oumar, et al.
Veröffentlicht: (2024)
ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation
von: Brassard, Ana, et al.
Veröffentlicht: (2024)
von: Brassard, Ana, et al.
Veröffentlicht: (2024)
Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki
von: Sasaki, Mutsumi, et al.
Veröffentlicht: (2025)
von: Sasaki, Mutsumi, et al.
Veröffentlicht: (2025)
Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
von: Dentan, Jérémie, et al.
Veröffentlicht: (2025)
von: Dentan, Jérémie, et al.
Veröffentlicht: (2025)
Repetition Neurons: How Do Language Models Produce Repetitions?
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
Scaling Laws for Fact Memorization of Large Language Models
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
von: Nwadike, Munachiso, et al.
Veröffentlicht: (2025)
von: Nwadike, Munachiso, et al.
Veröffentlicht: (2025)
Fact Recall, Heuristics or Pure Guesswork? Precise Interpretations of Language Models for Fact Completion
von: Saynova, Denitsa, et al.
Veröffentlicht: (2024)
von: Saynova, Denitsa, et al.
Veröffentlicht: (2024)
Memory Dial: A Training Framework for Controllable Memorization in Language Models
von: Zhang, Xiangbo, et al.
Veröffentlicht: (2026)
von: Zhang, Xiangbo, et al.
Veröffentlicht: (2026)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
von: Ye, Jiayuan, et al.
Veröffentlicht: (2026)
von: Ye, Jiayuan, et al.
Veröffentlicht: (2026)
Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models
von: Vladika, Juraj, et al.
Veröffentlicht: (2025)
von: Vladika, Juraj, et al.
Veröffentlicht: (2025)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
von: Doan, Nhi Hoai, et al.
Veröffentlicht: (2025)
von: Doan, Nhi Hoai, et al.
Veröffentlicht: (2025)
Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders
von: Inaba, Tatsuro, et al.
Veröffentlicht: (2025)
von: Inaba, Tatsuro, et al.
Veröffentlicht: (2025)
Reconsidering Degeneration of Token Embeddings with Definitions for Encoder-based Pre-trained Language Models
von: Zhang, Ying, et al.
Veröffentlicht: (2024)
von: Zhang, Ying, et al.
Veröffentlicht: (2024)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2025)
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2025)
Syntactic Learnability of Echo State Neural Language Models at Scale
von: Ueda, Ryo, et al.
Veröffentlicht: (2025)
von: Ueda, Ryo, et al.
Veröffentlicht: (2025)
Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models
von: Tang, Ethan
Veröffentlicht: (2026)
von: Tang, Ethan
Veröffentlicht: (2026)
Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
von: Hara, Tomomasa, et al.
Veröffentlicht: (2026)
von: Hara, Tomomasa, et al.
Veröffentlicht: (2026)
Memorization Dynamics in Knowledge Distillation for Language Models
von: Borkar, Jaydeep, et al.
Veröffentlicht: (2026)
von: Borkar, Jaydeep, et al.
Veröffentlicht: (2026)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2024)
Time Awareness in Large Language Models: Benchmarking Fact Recall Across Time
von: Herel, David, et al.
Veröffentlicht: (2024)
von: Herel, David, et al.
Veröffentlicht: (2024)
Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
Functional Abstraction of Knowledge Recall in Large Language Models
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
Tracing Relational Knowledge Recall in Large Language Models
von: Popovič, Nicholas, et al.
Veröffentlicht: (2026)
von: Popovič, Nicholas, et al.
Veröffentlicht: (2026)
Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
Large Language Models Are Human-Like Internally
von: Kuribayashi, Tatsuki, et al.
Veröffentlicht: (2025)
von: Kuribayashi, Tatsuki, et al.
Veröffentlicht: (2025)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
von: Pradeep, Ronak, et al.
Veröffentlicht: (2025)
von: Pradeep, Ronak, et al.
Veröffentlicht: (2025)
Quantifying Memorization and Detecting Training Data of Pre-trained Language Models using Japanese Newspaper
von: Ishihara, Shotaro, et al.
Veröffentlicht: (2024)
von: Ishihara, Shotaro, et al.
Veröffentlicht: (2024)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
von: Niwa, Ayana, et al.
Veröffentlicht: (2025)
von: Niwa, Ayana, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Monotonic Representation of Numeric Properties in Language Models
von: Heinzerling, Benjamin, et al.
Veröffentlicht: (2024) -
Representational Analysis of Binding in Language Models
von: Dai, Qin, et al.
Veröffentlicht: (2024) -
Cell-Based Representation of Relational Binding in Language Models
von: Dai, Qin, et al.
Veröffentlicht: (2026) -
Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference
von: Kamoda, Go, et al.
Veröffentlicht: (2025) -
TopK Language Models
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2025)