The Power of Power Law: Asymmetry Enables Compositional Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zixuan, Dang, Xingyu, Lee, Jason D., Lyu, Kaifeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910163943817216
author Wang, Zixuan
Dang, Xingyu
Lee, Jason D.
Lyu, Kaifeng
author_facet Wang, Zixuan
Dang, Xingyu
Lee, Jason D.
Lyu, Kaifeng
contents Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intuition suggests that reweighting or curating data towards a uniform distribution may help models better learn these long-tail skills, we find a counterintuitive result: across a wide range of compositional reasoning tasks, such as state tracking and multi-step arithmetic, training under power-law distributions consistently outperforms training under uniform distributions. To understand this advantage, we introduce a minimalist skill-composition task and show that learning under a power-law distribution provably requires significantly less training data. Our theoretical analysis reveals that power law sampling induces a beneficial asymmetry that improves the pathological loss landscape, which enables models to first acquire high-frequency skill compositions with low data complexity, which in turn serves as a stepping stone to efficiently learn rare long-tailed skills. Our results offer an alternative perspective on what constitutes an effective data distribution for training models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22951
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Power of Power Law: Asymmetry Enables Compositional Reasoning
Wang, Zixuan
Dang, Xingyu
Lee, Jason D.
Lyu, Kaifeng
Artificial Intelligence
Computation and Language
Machine Learning
Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intuition suggests that reweighting or curating data towards a uniform distribution may help models better learn these long-tail skills, we find a counterintuitive result: across a wide range of compositional reasoning tasks, such as state tracking and multi-step arithmetic, training under power-law distributions consistently outperforms training under uniform distributions. To understand this advantage, we introduce a minimalist skill-composition task and show that learning under a power-law distribution provably requires significantly less training data. Our theoretical analysis reveals that power law sampling induces a beneficial asymmetry that improves the pathological loss landscape, which enables models to first acquire high-frequency skill compositions with low data complexity, which in turn serves as a stepping stone to efficiently learn rare long-tailed skills. Our results offer an alternative perspective on what constitutes an effective data distribution for training models.
title The Power of Power Law: Asymmetry Enables Compositional Reasoning
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2604.22951