Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Smith, Matthew L., Shock, Jonathan P., Segun, Samuel T., Olatunji, Iyiola E., Bissyandé, Tegawendé F.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918509835976704
author Smith, Matthew L.
Shock, Jonathan P.
Segun, Samuel T.
Olatunji, Iyiola E.
Bissyandé, Tegawendé F.
author_facet Smith, Matthew L.
Shock, Jonathan P.
Segun, Samuel T.
Olatunji, Iyiola E.
Bissyandé, Tegawendé F.
contents While scaling laws govern aggregate large language model performance, no scaling law has linked factual recall to both model size and training-data composition. We evaluated 38 models on over 8,900 scholarly references evaluated by an automated reference verification system. Recall quality follows a sigmoid in the log-linear combination of model parameter count and topic representation in training data. These two variables alone explain 60% of the variance across 16 dense models from four families, rising to 74-94% within individual families. The form matches a superposition-inspired account in which recall is gated by a signal-to-noise ratio: signal strength scales with concept frequency and the noise floor with model capacity.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18732
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
Smith, Matthew L.
Shock, Jonathan P.
Segun, Samuel T.
Olatunji, Iyiola E.
Bissyandé, Tegawendé F.
Computation and Language
Artificial Intelligence
Machine Learning
While scaling laws govern aggregate large language model performance, no scaling law has linked factual recall to both model size and training-data composition. We evaluated 38 models on over 8,900 scholarly references evaluated by an automated reference verification system. Recall quality follows a sigmoid in the log-linear combination of model parameter count and topic representation in training data. These two variables alone explain 60% of the variance across 16 dense models from four families, rising to 74-94% within individual families. The form matches a superposition-inspired account in which recall is gated by a signal-to-noise ratio: signal strength scales with concept frequency and the noise floor with model capacity.
title Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.18732