On the Cost and Benefit of Chain of Thought: A Learning-Theoretic Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yue, Dong, Zhiyi, Cesari, Tommaso, Mao, Yongyi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911702203760640
author Zhang, Yue
Dong, Zhiyi
Cesari, Tommaso
Mao, Yongyi
author_facet Zhang, Yue
Dong, Zhiyi
Cesari, Tommaso
Mao, Yongyi
contents We develop a learning-theoretic framework for understanding Chain of Thought (CoT). We model CoT as the interaction between an answer map and a chain rule that generates intermediate questions autoregressively, and define the reasoning risk of a hypothesis under this interaction. Our first result is a tight canonical decomposition of this risk into two terms with opposing roles: an oracle-trajectory risk (OTR), which captures the benefit of CoT and reduces to a target-domain risk in a domain adaptation problem, and a trajectory-mismatch risk (TMR), which captures the cost of CoT through error accumulation along mismatched reasoning trajectories. We then show that this cost is unavoidable without structure: if any one of the loss, the hypothesis answer map, or the chain rule lacks stability, the TMR can be arbitrarily large even when the OTR is zero and the hypothesis is uniformly close to the ground truth. Conversely, under stability, we prove a tight upper bound on the TMR governed by an exact amplification factor that identifies bounded, linear, and exponential error-growth regimes. Together, these results give a precise theory of when CoT helps, when it hurts, and what controls the transition between the two.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21260
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the Cost and Benefit of Chain of Thought: A Learning-Theoretic Perspective
Zhang, Yue
Dong, Zhiyi
Cesari, Tommaso
Mao, Yongyi
Machine Learning
We develop a learning-theoretic framework for understanding Chain of Thought (CoT). We model CoT as the interaction between an answer map and a chain rule that generates intermediate questions autoregressively, and define the reasoning risk of a hypothesis under this interaction. Our first result is a tight canonical decomposition of this risk into two terms with opposing roles: an oracle-trajectory risk (OTR), which captures the benefit of CoT and reduces to a target-domain risk in a domain adaptation problem, and a trajectory-mismatch risk (TMR), which captures the cost of CoT through error accumulation along mismatched reasoning trajectories. We then show that this cost is unavoidable without structure: if any one of the loss, the hypothesis answer map, or the chain rule lacks stability, the TMR can be arbitrarily large even when the OTR is zero and the hypothesis is uniformly close to the ground truth. Conversely, under stability, we prove a tight upper bound on the TMR governed by an exact amplification factor that identifies bounded, linear, and exponential error-growth regimes. Together, these results give a precise theory of when CoT helps, when it hurts, and what controls the transition between the two.
title On the Cost and Benefit of Chain of Thought: A Learning-Theoretic Perspective
topic Machine Learning
url https://arxiv.org/abs/2605.21260