Universal One-third Time Scaling in Learning Peaked Distributions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Yizhou, Liu, Ziming, Pehlevan, Cengiz, Gore, Jeff
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916071268679680
author Liu, Yizhou
Liu, Ziming
Pehlevan, Cengiz
Gore, Jeff
author_facet Liu, Yizhou
Liu, Ziming
Pehlevan, Cengiz
Gore, Jeff
contents Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable. Through systematic analysis of toy models and empirical evaluation of LLMs, we show that this behavior can arise intrinsically from the use of softmax and cross-entropy. When learning peaked probability distributions, e.g., next-token distributions, these components generically yield power-law vanishing losses and gradients, regardless of many microscopic details, creating a fundamental optimization bottleneck. This ultimately leads to power-law time scaling of the loss with a universal exponent of $1/3$. Our results provide a mechanistic explanation for observed neural scaling and suggest new directions for improving LLM training efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03685
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Universal One-third Time Scaling in Learning Peaked Distributions
Liu, Yizhou
Liu, Ziming
Pehlevan, Cengiz
Gore, Jeff
Machine Learning
Artificial Intelligence
Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable. Through systematic analysis of toy models and empirical evaluation of LLMs, we show that this behavior can arise intrinsically from the use of softmax and cross-entropy. When learning peaked probability distributions, e.g., next-token distributions, these components generically yield power-law vanishing losses and gradients, regardless of many microscopic details, creating a fundamental optimization bottleneck. This ultimately leads to power-law time scaling of the loss with a universal exponent of $1/3$. Our results provide a mechanistic explanation for observed neural scaling and suggest new directions for improving LLM training efficiency.
title Universal One-third Time Scaling in Learning Peaked Distributions
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.03685