Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Ivan, Berg-Kirkpatrick, Taylor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909849519915008
author Lee, Ivan
Berg-Kirkpatrick, Taylor
author_facet Lee, Ivan
Berg-Kirkpatrick, Taylor
contents Recent studies suggest that very small language models (SLMs) can generate surprisingly coherent text when trained on simplified, child-directed corpora such as TinyStories. These findings have been interpreted as evidence that readability -- characterized by accessible vocabulary, familiar narrative structure, and simple syntax -- plays a key role in enabling such capabilities to emerge. In this paper, we challenge that interpretation. We construct synthetic datasets with matched structure but varied readability, and find that readability alone does not predict coherence or learning efficiency in SLMs. Models trained on complex, adult-level text perform comparably to those trained on simplified language, and even exhibit faster development of coherence during training. Instead, we show that statistical simplicity, as measured by n-gram diversity, is a stronger predictor of learnability. Our findings caution against the growing trend of anthropomorphizing language model training -- drawing parallels to human cognitive development without empirical basis -- and argue for more precise reasoning about what properties actually support capability emergence in small models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
Lee, Ivan
Berg-Kirkpatrick, Taylor
Computation and Language
Artificial Intelligence
Machine Learning
Recent studies suggest that very small language models (SLMs) can generate surprisingly coherent text when trained on simplified, child-directed corpora such as TinyStories. These findings have been interpreted as evidence that readability -- characterized by accessible vocabulary, familiar narrative structure, and simple syntax -- plays a key role in enabling such capabilities to emerge. In this paper, we challenge that interpretation. We construct synthetic datasets with matched structure but varied readability, and find that readability alone does not predict coherence or learning efficiency in SLMs. Models trained on complex, adult-level text perform comparably to those trained on simplified language, and even exhibit faster development of coherence during training. Instead, we show that statistical simplicity, as measured by n-gram diversity, is a stronger predictor of learnability. Our findings caution against the growing trend of anthropomorphizing language model training -- drawing parallels to human cognitive development without empirical basis -- and argue for more precise reasoning about what properties actually support capability emergence in small models.
title Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.13915