LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gottesman, Daniela, Gilae-Dotan, Alon, Cohen, Ido, Gur-Arieh, Yoav, Mosbach, Marius, Yoran, Ori, Geva, Mor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915477881618432
author Gottesman, Daniela
Gilae-Dotan, Alon
Cohen, Ido
Gur-Arieh, Yoav
Mosbach, Marius
Yoran, Ori
Geva, Mor
author_facet Gottesman, Daniela
Gilae-Dotan, Alon
Cohen, Ido
Gur-Arieh, Yoav
Mosbach, Marius
Yoran, Ori
Geva, Mor
contents Language models (LMs) increasingly drive real-world applications that require world knowledge. However, the internal processes through which models turn data into representations of knowledge and beliefs about the world, are poorly understood. Insights into these processes could pave the way for developing LMs with knowledge representations that are more consistent, robust, and complete. To facilitate studying these questions, we present LMEnt, a suite for analyzing knowledge acquisition in LMs during pretraining. LMEnt introduces: (1) a knowledge-rich pretraining corpus, fully annotated with entity mentions, based on Wikipedia, (2) an entity-based retrieval method over pretraining data that outperforms previous approaches by as much as 80.4%, and (3) 12 pretrained models with up to 1B parameters and 4K intermediate checkpoints, with comparable performance to popular open-sourced models on knowledge benchmarks. Together, these resources provide a controlled environment for analyzing connections between entity mentions in pretraining and downstream performance, and the effects of causal interventions in pretraining data. We show the utility of LMEnt by studying knowledge acquisition across checkpoints, finding that fact frequency is key, but does not fully explain learning trends. We release LMEnt to support studies of knowledge in LMs, including knowledge representations, plasticity, editing, attribution, and learning dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
Gottesman, Daniela
Gilae-Dotan, Alon
Cohen, Ido
Gur-Arieh, Yoav
Mosbach, Marius
Yoran, Ori
Geva, Mor
Computation and Language
Language models (LMs) increasingly drive real-world applications that require world knowledge. However, the internal processes through which models turn data into representations of knowledge and beliefs about the world, are poorly understood. Insights into these processes could pave the way for developing LMs with knowledge representations that are more consistent, robust, and complete. To facilitate studying these questions, we present LMEnt, a suite for analyzing knowledge acquisition in LMs during pretraining. LMEnt introduces: (1) a knowledge-rich pretraining corpus, fully annotated with entity mentions, based on Wikipedia, (2) an entity-based retrieval method over pretraining data that outperforms previous approaches by as much as 80.4%, and (3) 12 pretrained models with up to 1B parameters and 4K intermediate checkpoints, with comparable performance to popular open-sourced models on knowledge benchmarks. Together, these resources provide a controlled environment for analyzing connections between entity mentions in pretraining and downstream performance, and the effects of causal interventions in pretraining data. We show the utility of LMEnt by studying knowledge acquisition across checkpoints, finding that fact frequency is key, but does not fully explain learning trends. We release LMEnt to support studies of knowledge in LMs, including knowledge representations, plasticity, editing, attribution, and learning dynamics.
title LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
topic Computation and Language
url https://arxiv.org/abs/2509.03405