Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schumann, Raphael, Riezler, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911185758060544
author Schumann, Raphael
Riezler, Stefan
author_facet Schumann, Raphael
Riezler, Stefan
contents Reasoning quality in large language models depends not only on producing correct answers but also on generating valid intermediate steps. We study this through multiple-choice question answering (MCQA), which provides a controlled setting with fixed answer options. Our analysis shows that when questions are effectively unsolvable for a model, spurious chains of thought (CoTs) are more likely to appear, leading to false positives. By estimating the solvability of each question, we uncover an intermediate regime where learning is most effective. Building on this insight, we adapt outcome-supervised reward models and reinforcement learning with group-relative advantage to incorporate solvability into their objectives. Across experiments on math and multimodal datasets, these modifications consistently yield higher rates of process-correct reasoning and, in reinforcement learning, improved answer accuracy as well. Our results highlight solvability as a key factor for reducing hallucinations and increasing reliability in CoT reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25941
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
Schumann, Raphael
Riezler, Stefan
Artificial Intelligence
Computation and Language
Reasoning quality in large language models depends not only on producing correct answers but also on generating valid intermediate steps. We study this through multiple-choice question answering (MCQA), which provides a controlled setting with fixed answer options. Our analysis shows that when questions are effectively unsolvable for a model, spurious chains of thought (CoTs) are more likely to appear, leading to false positives. By estimating the solvability of each question, we uncover an intermediate regime where learning is most effective. Building on this insight, we adapt outcome-supervised reward models and reinforcement learning with group-relative advantage to incorporate solvability into their objectives. Across experiments on math and multimodal datasets, these modifications consistently yield higher rates of process-correct reasoning and, in reinforcement learning, improved answer accuracy as well. Our results highlight solvability as a key factor for reducing hallucinations and increasing reliability in CoT reasoning.
title Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.25941