Pre-training Limited Memory Language Models with Internal and External Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Linxi, Zalouk, Sofian, Belardi, Christian K., Lovelace, Justin, Zhou, Jin Peng, Noonan, Ryan Thomas, Go, Dongyoung, Weinberger, Kilian Q., Artzi, Yoav, Sun, Jennifer J.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912624756654080
author Zhao, Linxi
Zalouk, Sofian
Belardi, Christian K.
Lovelace, Justin
Zhou, Jin Peng
Noonan, Ryan Thomas
Go, Dongyoung
Weinberger, Kilian Q.
Artzi, Yoav
Sun, Jennifer J.
author_facet Zhao, Linxi
Zalouk, Sofian
Belardi, Christian K.
Lovelace, Justin
Zhou, Jin Peng
Noonan, Ryan Thomas
Go, Dongyoung
Weinberger, Kilian Q.
Artzi, Yoav
Sun, Jennifer J.
contents Neural language models are black-boxes--both linguistic patterns and factual knowledge are distributed across billions of opaque parameters. This entangled encoding makes it difficult to reliably inspect, verify, or update specific facts. We introduce Limited Memory Language Models (LMLM), a new class of language models that externalizes factual knowledge to external database during pre-training rather than memorizing them. Our pre-training approach strategically masks externally retrieved factual values from the training loss, thereby teaching the model to perform targeted lookups rather than relying on memorization in model weights. Our experiments demonstrate that LMLMs achieve competitive performance compared to significantly larger LLMs on standard benchmarks, while offering the advantages of explicit, editable, and verifiable knowledge bases.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15962
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pre-training Limited Memory Language Models with Internal and External Knowledge
Zhao, Linxi
Zalouk, Sofian
Belardi, Christian K.
Lovelace, Justin
Zhou, Jin Peng
Noonan, Ryan Thomas
Go, Dongyoung
Weinberger, Kilian Q.
Artzi, Yoav
Sun, Jennifer J.
Computation and Language
Artificial Intelligence
Machine Learning
Neural language models are black-boxes--both linguistic patterns and factual knowledge are distributed across billions of opaque parameters. This entangled encoding makes it difficult to reliably inspect, verify, or update specific facts. We introduce Limited Memory Language Models (LMLM), a new class of language models that externalizes factual knowledge to external database during pre-training rather than memorizing them. Our pre-training approach strategically masks externally retrieved factual values from the training loss, thereby teaching the model to perform targeted lookups rather than relying on memorization in model weights. Our experiments demonstrate that LMLMs achieve competitive performance compared to significantly larger LLMs on standard benchmarks, while offering the advantages of explicit, editable, and verifiable knowledge bases.
title Pre-training Limited Memory Language Models with Internal and External Knowledge
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.15962