Correct Reasoning Paths Visit Shared Decision Pivots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cho, Dongkyu, Zhang, Amy B. Z., Fehri, Bilel, Wang, Sheng, Chunara, Rumi, Cai, Hengrui, Song, Rui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910015068045312
author Cho, Dongkyu
Zhang, Amy B. Z.
Fehri, Bilel
Wang, Sheng
Chunara, Rumi
Cai, Hengrui
Song, Rui
author_facet Cho, Dongkyu
Zhang, Amy B. Z.
Fehri, Bilel
Wang, Sheng
Chunara, Rumi
Cai, Hengrui
Song, Rui
contents Chain-of-thought (CoT) reasoning exposes the intermediate thinking process of large language models (LLMs), yet verifying those traces at scale remains unsolved. In response, we introduce the idea of decision pivots-minimal, verifiable checkpoints that any correct reasoning path must visit. We hypothesize that correct reasoning, though stylistically diverse, converge on the same pivot set, while incorrect ones violate at least one pivot. Leveraging this property, we propose a self-training pipeline that (i) samples diverse reasoning paths and mines shared decision pivots, (ii) compresses each trace into pivot-focused short-path reasoning using an auxiliary verifier, and (iii) post-trains the model using its self-generated outputs. The proposed method aligns reasoning without ground truth reasoning data or external metrics. Experiments on standard benchmarks such as LogiQA, MedQA, and MATH500 show the effectiveness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Correct Reasoning Paths Visit Shared Decision Pivots
Cho, Dongkyu
Zhang, Amy B. Z.
Fehri, Bilel
Wang, Sheng
Chunara, Rumi
Cai, Hengrui
Song, Rui
Artificial Intelligence
Chain-of-thought (CoT) reasoning exposes the intermediate thinking process of large language models (LLMs), yet verifying those traces at scale remains unsolved. In response, we introduce the idea of decision pivots-minimal, verifiable checkpoints that any correct reasoning path must visit. We hypothesize that correct reasoning, though stylistically diverse, converge on the same pivot set, while incorrect ones violate at least one pivot. Leveraging this property, we propose a self-training pipeline that (i) samples diverse reasoning paths and mines shared decision pivots, (ii) compresses each trace into pivot-focused short-path reasoning using an auxiliary verifier, and (iii) post-trains the model using its self-generated outputs. The proposed method aligns reasoning without ground truth reasoning data or external metrics. Experiments on standard benchmarks such as LogiQA, MedQA, and MATH500 show the effectiveness of our method.
title Correct Reasoning Paths Visit Shared Decision Pivots
topic Artificial Intelligence
url https://arxiv.org/abs/2509.21549