Correct Reasoning Paths Visit Shared Decision Pivots
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910015068045312 |
|---|---|
| author | Cho, Dongkyu Zhang, Amy B. Z. Fehri, Bilel Wang, Sheng Chunara, Rumi Cai, Hengrui Song, Rui |
| author_facet | Cho, Dongkyu Zhang, Amy B. Z. Fehri, Bilel Wang, Sheng Chunara, Rumi Cai, Hengrui Song, Rui |
| contents | Chain-of-thought (CoT) reasoning exposes the intermediate thinking process of large language models (LLMs), yet verifying those traces at scale remains unsolved. In response, we introduce the idea of decision pivots-minimal, verifiable checkpoints that any correct reasoning path must visit. We hypothesize that correct reasoning, though stylistically diverse, converge on the same pivot set, while incorrect ones violate at least one pivot. Leveraging this property, we propose a self-training pipeline that (i) samples diverse reasoning paths and mines shared decision pivots, (ii) compresses each trace into pivot-focused short-path reasoning using an auxiliary verifier, and (iii) post-trains the model using its self-generated outputs. The proposed method aligns reasoning without ground truth reasoning data or external metrics. Experiments on standard benchmarks such as LogiQA, MedQA, and MATH500 show the effectiveness of our method. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_21549 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Correct Reasoning Paths Visit Shared Decision Pivots Cho, Dongkyu Zhang, Amy B. Z. Fehri, Bilel Wang, Sheng Chunara, Rumi Cai, Hengrui Song, Rui Artificial Intelligence Chain-of-thought (CoT) reasoning exposes the intermediate thinking process of large language models (LLMs), yet verifying those traces at scale remains unsolved. In response, we introduce the idea of decision pivots-minimal, verifiable checkpoints that any correct reasoning path must visit. We hypothesize that correct reasoning, though stylistically diverse, converge on the same pivot set, while incorrect ones violate at least one pivot. Leveraging this property, we propose a self-training pipeline that (i) samples diverse reasoning paths and mines shared decision pivots, (ii) compresses each trace into pivot-focused short-path reasoning using an auxiliary verifier, and (iii) post-trains the model using its self-generated outputs. The proposed method aligns reasoning without ground truth reasoning data or external metrics. Experiments on standard benchmarks such as LogiQA, MedQA, and MATH500 show the effectiveness of our method. |
| title | Correct Reasoning Paths Visit Shared Decision Pivots |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2509.21549 |