How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Minsung, Kim, Dong-Kyum, Kwon, Jea, Yang, Nakyeong, Jung, Kyomin, Cha, Meeyoung
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918453875572736
author Kim, Minsung
Kim, Dong-Kyum
Kwon, Jea
Yang, Nakyeong
Jung, Kyomin
Cha, Meeyoung
author_facet Kim, Minsung
Kim, Dong-Kyum
Kwon, Jea
Yang, Nakyeong
Jung, Kyomin
Cha, Meeyoung
contents Large language models leverage both parametric knowledge acquired during pretraining and in-context knowledge provided at inference time. Crucially, when these sources conflict, models arbitrate based on their internal confidence, preferring parametric knowledge for high-confidence facts while deferring to context for less familiar ones. However, the training conditions that give rise to these fundamental behaviors remain unclear. Here we conduct controlled experiments using synthetic corpora to identify the specific data properties that shape knowledge utilization. Our results reveal a counterintuitive finding: the robust, balanced use of both knowledge sources is an emergent property that requires the co-occurrence of three factors typically considered detrimental, including (i) intra-document repetition, (ii) a moderate degree of intra-document inconsistency, and (iii) a skewed knowledge distribution. We further show that these dynamics arise in real-world language model pretraining and analyze how post-training procedures reshape arbitration strategies. Together, our findings provide empirical guidance for designing training data that supports the reliable integration of parametric and in-context knowledge in language models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02370
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
Kim, Minsung
Kim, Dong-Kyum
Kwon, Jea
Yang, Nakyeong
Jung, Kyomin
Cha, Meeyoung
Computation and Language
Artificial Intelligence
Large language models leverage both parametric knowledge acquired during pretraining and in-context knowledge provided at inference time. Crucially, when these sources conflict, models arbitrate based on their internal confidence, preferring parametric knowledge for high-confidence facts while deferring to context for less familiar ones. However, the training conditions that give rise to these fundamental behaviors remain unclear. Here we conduct controlled experiments using synthetic corpora to identify the specific data properties that shape knowledge utilization. Our results reveal a counterintuitive finding: the robust, balanced use of both knowledge sources is an emergent property that requires the co-occurrence of three factors typically considered detrimental, including (i) intra-document repetition, (ii) a moderate degree of intra-document inconsistency, and (iii) a skewed knowledge distribution. We further show that these dynamics arise in real-world language model pretraining and analyze how post-training procedures reshape arbitration strategies. Together, our findings provide empirical guidance for designing training data that supports the reliable integration of parametric and in-context knowledge in language models.
title How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.02370