A Primer in Post-Training Reasoning Data: What We Know About How It Works

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Yaoming, Zhao, Guangxiang, Shi, Qilong, Sun, Lin, Zhang, Xiangzheng, Yang, Tong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911741337665536
author Li, Yaoming
Zhao, Guangxiang
Shi, Qilong
Sun, Lin
Zhang, Xiangzheng
Yang, Tong
author_facet Li, Yaoming
Zhao, Guangxiang
Shi, Qilong
Sun, Lin
Zhang, Xiangzheng
Yang, Tong
contents Post-training has become a primary driver of recent progress in large reasoning models, and reasoning data are often the key variable determining whether this stage succeeds. Work on post-training reasoning data has grown rapidly, yet this literature remains scattered across dataset papers, reinforcement-learning recipes, reward-model studies, benchmarks, and frontier system reports. This paper is the first primer to synthesize over 150 key public studies and system reports on post-training reasoning data. We organize the field around four questions: what data objects exist, what makes them useful, how they are constructed, and how they scale. Together, this organization provides an attribution framework for future reasoning-data releases and post-training recipes.
format Preprint
id arxiv_https___arxiv_org_abs_2606_02113
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Primer in Post-Training Reasoning Data: What We Know About How It Works
Li, Yaoming
Zhao, Guangxiang
Shi, Qilong
Sun, Lin
Zhang, Xiangzheng
Yang, Tong
Computation and Language
Artificial Intelligence
Post-training has become a primary driver of recent progress in large reasoning models, and reasoning data are often the key variable determining whether this stage succeeds. Work on post-training reasoning data has grown rapidly, yet this literature remains scattered across dataset papers, reinforcement-learning recipes, reward-model studies, benchmarks, and frontier system reports. This paper is the first primer to synthesize over 150 key public studies and system reports on post-training reasoning data. We organize the field around four questions: what data objects exist, what makes them useful, how they are constructed, and how they scale. Together, this organization provides an attribution framework for future reasoning-data releases and post-training recipes.
title A Primer in Post-Training Reasoning Data: What We Know About How It Works
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2606.02113