Systematic Generalization in Language Models Scales with Information Entropy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wold, Sondre, Charpentier, Lucas Georges Gabriel, Simon, Étienne
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909624501796864
author Wold, Sondre
Charpentier, Lucas Georges Gabriel
Simon, Étienne
author_facet Wold, Sondre
Charpentier, Lucas Georges Gabriel
Simon, Étienne
contents Systematic generalization remains challenging for current language models, which are known to be both sensitive to semantically similar permutations of the input and to struggle with known concepts presented in novel contexts. Although benchmarks exist for assessing compositional behavior, it is unclear how to measure the difficulty of a systematic generalization problem. In this work, we show how one aspect of systematic generalization can be described by the entropy of the distribution of component parts in the training data. We formalize a framework for measuring entropy in a sequence-to-sequence task and find that the performance of popular model architectures scales with the entropy. Our work connects systematic generalization to information efficiency, and our results indicate that success at high entropy can be achieved even without built-in priors, and that success at low entropy can serve as a target for assessing progress towards robust systematic generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13089
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Systematic Generalization in Language Models Scales with Information Entropy
Wold, Sondre
Charpentier, Lucas Georges Gabriel
Simon, Étienne
Computation and Language
Systematic generalization remains challenging for current language models, which are known to be both sensitive to semantically similar permutations of the input and to struggle with known concepts presented in novel contexts. Although benchmarks exist for assessing compositional behavior, it is unclear how to measure the difficulty of a systematic generalization problem. In this work, we show how one aspect of systematic generalization can be described by the entropy of the distribution of component parts in the training data. We formalize a framework for measuring entropy in a sequence-to-sequence task and find that the performance of popular model architectures scales with the entropy. Our work connects systematic generalization to information efficiency, and our results indicate that success at high entropy can be achieved even without built-in priors, and that success at low entropy can serve as a target for assessing progress towards robust systematic generalization.
title Systematic Generalization in Language Models Scales with Information Entropy
topic Computation and Language
url https://arxiv.org/abs/2505.13089