The Frequency Confound in Language-Model Surprisal and Metaphor Novelty

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Momen, Omar, Zarrieß, Sina
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910227877593088
author Momen, Omar
Zarrieß, Sina
author_facet Momen, Omar
Zarrieß, Sina
contents Language-model (LM) surprisal is widely used as a proxy for contextual predictability and has been reported to correlate with metaphor novelty judgments. However, surprisal is tightly intertwined with lexical frequency. We explore this interaction on metaphor novelty ratings using two different word frequency measures. We analyse surprisal estimates from eight Pythia model sizes and 154 training checkpoints. Across settings, word frequency is a stronger predictor of metaphor novelty than surprisal. Across training stages, the surprisal--novelty association peaks at an early stage and then falls again, mirroring a similarly timed increase in the surprisal--frequency association. These results suggest that the often-reported optimal LM surprisal settings may incorrectly associate contextual predictability with metaphor novelty and processing difficulty, whereas lexical frequency may be the major underlying factor.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06506
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
Momen, Omar
Zarrieß, Sina
Computation and Language
Language-model (LM) surprisal is widely used as a proxy for contextual predictability and has been reported to correlate with metaphor novelty judgments. However, surprisal is tightly intertwined with lexical frequency. We explore this interaction on metaphor novelty ratings using two different word frequency measures. We analyse surprisal estimates from eight Pythia model sizes and 154 training checkpoints. Across settings, word frequency is a stronger predictor of metaphor novelty than surprisal. Across training stages, the surprisal--novelty association peaks at an early stage and then falls again, mirroring a similarly timed increase in the surprisal--frequency association. These results suggest that the often-reported optimal LM surprisal settings may incorrectly associate contextual predictability with metaphor novelty and processing difficulty, whereas lexical frequency may be the major underlying factor.
title The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
topic Computation and Language
url https://arxiv.org/abs/2605.06506