Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hakimi, Ahmad Dawar, Modarressi, Ali, Wicke, Philipp, Schütze, Hinrich
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913874834358272
author Hakimi, Ahmad Dawar
Modarressi, Ali
Wicke, Philipp
Schütze, Hinrich
author_facet Hakimi, Ahmad Dawar
Modarressi, Ali
Wicke, Philipp
Schütze, Hinrich
contents Understanding how large language models (LLMs) acquire and store factual knowledge is crucial for enhancing their interpretability and reliability. In this work, we analyze the evolution of factual knowledge representation in the OLMo-7B model by tracking the roles of its attention heads and feed forward networks (FFNs) over the course of pre-training. We classify these components into four roles: general, entity, relation-answer, and fact-answer specific, and examine their stability and transitions. Our results show that LLMs initially depend on broad, general-purpose components, which later specialize as training progresses. Once the model reliably predicts answers, some components are repurposed, suggesting an adaptive learning process. Notably, attention heads display the highest turnover. We also present evidence that FFNs remain more stable throughout training. Furthermore, our probing experiments reveal that location-based relations converge to high accuracy earlier in training than name-based relations, highlighting how task complexity shapes acquisition dynamics. These insights offer a mechanistic view of knowledge formation in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03434
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
Hakimi, Ahmad Dawar
Modarressi, Ali
Wicke, Philipp
Schütze, Hinrich
Computation and Language
Understanding how large language models (LLMs) acquire and store factual knowledge is crucial for enhancing their interpretability and reliability. In this work, we analyze the evolution of factual knowledge representation in the OLMo-7B model by tracking the roles of its attention heads and feed forward networks (FFNs) over the course of pre-training. We classify these components into four roles: general, entity, relation-answer, and fact-answer specific, and examine their stability and transitions. Our results show that LLMs initially depend on broad, general-purpose components, which later specialize as training progresses. Once the model reliably predicts answers, some components are repurposed, suggesting an adaptive learning process. Notably, attention heads display the highest turnover. We also present evidence that FFNs remain more stable throughout training. Furthermore, our probing experiments reveal that location-based relations converge to high accuracy earlier in training than name-based relations, highlighting how task complexity shapes acquisition dynamics. These insights offer a mechanistic view of knowledge formation in LLMs.
title Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2506.03434