To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Karan, Yu, Michael, Gangal, Varun, Tao, Zhuofu, Kumar, Sachin, Liu, Emmy, Feng, Steven Y.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917377157890048
author Singh, Karan
Yu, Michael
Gangal, Varun
Tao, Zhuofu
Kumar, Sachin
Liu, Emmy
Feng, Steven Y.
author_facet Singh, Karan
Yu, Michael
Gangal, Varun
Tao, Zhuofu
Kumar, Sachin
Liu, Emmy
Feng, Steven Y.
contents Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-intensive situations. However, the relationship between parametric knowledge acquired during pretraining and non-parametric knowledge accessed via retrieval remains poorly understood, especially under fixed data budgets. In this work, we systematically study the trade-off between pretraining corpus size and retrieval store size across a wide range of model and data scales. We train OLMo-2-based LMs ranging from 30M to 3B parameters on up to 100B tokens of DCLM data, while varying both pretraining data scale (1-150x the number of parameters) and retrieval store size (1-20x), and evaluate performance across a diverse suite of benchmarks spanning reasoning, scientific QA, and open-domain QA. We find that retrieval consistently improves performance over parametric-only baselines across model scales and introduce a three-dimensional scaling framework that models performance as a function of model size, pretraining tokens, and retrieval corpus size. This scaling manifold enables us to estimate optimal allocations of a fixed data budget between pretraining and retrieval, revealing that the marginal utility of retrieval depends strongly on model scale, task type, and the degree of pretraining saturation. Our results provide a quantitative foundation for understanding when and how retrieval should complement pretraining, offering practical guidance for allocating data resources in the design of scalable language modeling systems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00715
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
Singh, Karan
Yu, Michael
Gangal, Varun
Tao, Zhuofu
Kumar, Sachin
Liu, Emmy
Feng, Steven Y.
Computation and Language
Artificial Intelligence
Machine Learning
Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-intensive situations. However, the relationship between parametric knowledge acquired during pretraining and non-parametric knowledge accessed via retrieval remains poorly understood, especially under fixed data budgets. In this work, we systematically study the trade-off between pretraining corpus size and retrieval store size across a wide range of model and data scales. We train OLMo-2-based LMs ranging from 30M to 3B parameters on up to 100B tokens of DCLM data, while varying both pretraining data scale (1-150x the number of parameters) and retrieval store size (1-20x), and evaluate performance across a diverse suite of benchmarks spanning reasoning, scientific QA, and open-domain QA. We find that retrieval consistently improves performance over parametric-only baselines across model scales and introduce a three-dimensional scaling framework that models performance as a function of model size, pretraining tokens, and retrieval corpus size. This scaling manifold enables us to estimate optimal allocations of a fixed data budget between pretraining and retrieval, revealing that the marginal utility of retrieval depends strongly on model scale, task type, and the degree of pretraining saturation. Our results provide a quantitative foundation for understanding when and how retrieval should complement pretraining, offering practical guidance for allocating data resources in the design of scalable language modeling systems.
title To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.00715