The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pengmei, Zihan, Mavromatis, Costas, Shen, Zhengyuan, Zhang, Yunyi, Ioannidis, Vassilis N., Rangwala, Huzefa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911240210612224
author Pengmei, Zihan
Mavromatis, Costas
Shen, Zhengyuan
Zhang, Yunyi
Ioannidis, Vassilis N.
Rangwala, Huzefa
author_facet Pengmei, Zihan
Mavromatis, Costas
Shen, Zhengyuan
Zhang, Yunyi
Ioannidis, Vassilis N.
Rangwala, Huzefa
contents Chain-of-thought (CoT) supervision can substantially improve transformer performance, yet the mechanisms by which models learn to follow and benefit from CoT remain poorly understood. We investigate these learning dynamics through the lens of grokking by pretraining transformers on symbolic reasoning tasks with tunable algorithmic complexity and controllable data composition to study their generalization. Models were trained under two settings: (i) producing only final answers, and (ii) emitting explicit CoT traces before answering. Our results show that while CoT generally improves task performance, its benefits depend on task complexity. To quantify these effects, we model the accuracy of the logarithmic training steps with a three-parameter logistic curve, revealing how the learning speed and shape vary with task complexity, data distribution, and the presence of CoT supervision. We also uncover a transient trace unfaithfulness phase: early in training, models often produce correct answers while skipping or contradicting CoT steps, before later aligning their reasoning traces with answers. Empirically, we (1) demonstrate that CoT accelerates generalization but does not overcome tasks with higher algorithmic complexity, such as finding list intersections; (2) introduce a kinetic modeling framework for understanding transformer learning; (3) characterize trace faithfulness as a dynamic property that emerges over training; and (4) show CoT alters internal transformer computation mechanistically.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25791
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
Pengmei, Zihan
Mavromatis, Costas
Shen, Zhengyuan
Zhang, Yunyi
Ioannidis, Vassilis N.
Rangwala, Huzefa
Machine Learning
Artificial Intelligence
Chain-of-thought (CoT) supervision can substantially improve transformer performance, yet the mechanisms by which models learn to follow and benefit from CoT remain poorly understood. We investigate these learning dynamics through the lens of grokking by pretraining transformers on symbolic reasoning tasks with tunable algorithmic complexity and controllable data composition to study their generalization. Models were trained under two settings: (i) producing only final answers, and (ii) emitting explicit CoT traces before answering. Our results show that while CoT generally improves task performance, its benefits depend on task complexity. To quantify these effects, we model the accuracy of the logarithmic training steps with a three-parameter logistic curve, revealing how the learning speed and shape vary with task complexity, data distribution, and the presence of CoT supervision. We also uncover a transient trace unfaithfulness phase: early in training, models often produce correct answers while skipping or contradicting CoT steps, before later aligning their reasoning traces with answers. Empirically, we (1) demonstrate that CoT accelerates generalization but does not overcome tasks with higher algorithmic complexity, such as finding list intersections; (2) introduce a kinetic modeling framework for understanding transformer learning; (3) characterize trace faithfulness as a dynamic property that emerges over training; and (4) show CoT alters internal transformer computation mechanistically.
title The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.25791