The Scaling Hypothesis Is Language-Contingent: Evidence from Cross-Linguistic Training Dynamics

Fuente: Zenodo
Guardado en:
Detalles Bibliográficos
Autor principal: Wasserman, Adam Zachary
Formato: Recurso digital
Lenguaje:inglés
Publicado: Zenodo 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866902182363660288
author Wasserman, Adam Zachary
author_facet Wasserman, Adam Zachary
contents <p>The scaling hypothesis holds that large language model performance improves predictably with increased compute, data, and parameters, following power-law relationships assumed to be universal [<span>Kaplan et al.</span>, <span>2020</span>, <span>Hoffmann et al.</span>, <span>2022</span>]. We test this assumption via a pre-registered controlled ablation (Pre-registration: OSF 10.17605/OSF.IO/SJ48B; Project: OSF 10.17605/OSF.IO/2PG8S), training identical 125M-parameter transformers on matched English and French corpora from C4, holding all hyperparameters constant. Confirming our pre-registered prediction, we observe dramatically divergent learning trajectories: French achieves grammatical competence (100% on agreement probes) at 197M tokens and maintains it through experiment completion at 181K steps (∼3B tokens), while English remains at chance level (40%) throughout, a >15x difference in emergence threshold. Perplexity trajectories show French approaching near-final values (PPL∼27) while English remains elevated (PPL∼1340), a 50x ratio at matched training steps. Cross-study comparison with Pythia 125M [<span>Biderman et al.</span>, <span>2023</span>], which required∼300B tokens to reach comparable perplexity, serves two functions: it validates that our English model performs as expected (consistent with established scaling behavior), and it suggests French may be 50–100x more training-efficient than English. These results support our hypothesis that morphologically rich languages provide redundant grammatical signals that accelerate structural learning. Critically, we show that perplexity and grammatical accuracy are orthogonal dimensions governed by different determinants: distributional coherence and morphological explicitness, respectively. This explains why English models can improve perplexity indefinitely while never acquiring grammar—standard evaluation metrics miss structural learning deficits entirely. The scaling hypothesis is language-contingent, not universal.</p> <p>Note: Pre-registered 350M experiments are complete but inconclusive due to batch size constraints. French 350M reached only 70% accuracy after 819M tokens (4×the tokens French 125M needed to emerge), suggesting scale may be counterproductive for morphologically rich languages. English 350M remained at 40% accuracy. We are re-running 350M experiments to 3.3B tokens to match the 125M token budget and will publish updated results. Training logs: <span>https://github.com/ </span>adamzwasserman/fractal-language</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19423151
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle The Scaling Hypothesis Is Language-Contingent: Evidence from Cross-Linguistic Training Dynamics
Wasserman, Adam Zachary
scaling laws
morphology
cross-linguistic
emergence
training dynamics
<p>The scaling hypothesis holds that large language model performance improves predictably with increased compute, data, and parameters, following power-law relationships assumed to be universal [<span>Kaplan et al.</span>, <span>2020</span>, <span>Hoffmann et al.</span>, <span>2022</span>]. We test this assumption via a pre-registered controlled ablation (Pre-registration: OSF 10.17605/OSF.IO/SJ48B; Project: OSF 10.17605/OSF.IO/2PG8S), training identical 125M-parameter transformers on matched English and French corpora from C4, holding all hyperparameters constant. Confirming our pre-registered prediction, we observe dramatically divergent learning trajectories: French achieves grammatical competence (100% on agreement probes) at 197M tokens and maintains it through experiment completion at 181K steps (∼3B tokens), while English remains at chance level (40%) throughout, a >15x difference in emergence threshold. Perplexity trajectories show French approaching near-final values (PPL∼27) while English remains elevated (PPL∼1340), a 50x ratio at matched training steps. Cross-study comparison with Pythia 125M [<span>Biderman et al.</span>, <span>2023</span>], which required∼300B tokens to reach comparable perplexity, serves two functions: it validates that our English model performs as expected (consistent with established scaling behavior), and it suggests French may be 50–100x more training-efficient than English. These results support our hypothesis that morphologically rich languages provide redundant grammatical signals that accelerate structural learning. Critically, we show that perplexity and grammatical accuracy are orthogonal dimensions governed by different determinants: distributional coherence and morphological explicitness, respectively. This explains why English models can improve perplexity indefinitely while never acquiring grammar—standard evaluation metrics miss structural learning deficits entirely. The scaling hypothesis is language-contingent, not universal.</p> <p>Note: Pre-registered 350M experiments are complete but inconclusive due to batch size constraints. French 350M reached only 70% accuracy after 819M tokens (4×the tokens French 125M needed to emerge), suggesting scale may be counterproductive for morphologically rich languages. English 350M remained at 40% accuracy. We are re-running 350M experiments to 3.3B tokens to match the 125M token budget and will publish updated results. Training logs: <span>https://github.com/ </span>adamzwasserman/fractal-language</p>
title The Scaling Hypothesis Is Language-Contingent: Evidence from Cross-Linguistic Training Dynamics
topic scaling laws
morphology
cross-linguistic
emergence
training dynamics
url https://doi.org/10.5281/zenodo.19423151