Improving Rationality in the Reasoning Process of Language Models through Self-playing Game

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Pinzheng, Li, Juntao, Tang, Zecheng, Gui, Haijia, zhang, Min
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916828691824640
author Wang, Pinzheng
Li, Juntao
Tang, Zecheng
Gui, Haijia
zhang, Min
author_facet Wang, Pinzheng
Li, Juntao
Tang, Zecheng
Gui, Haijia
zhang, Min
contents Large language models (LLMs) have demonstrated considerable reasoning abilities in various tasks such as mathematics and coding. However, recent studies indicate that even the best models lack true comprehension of their reasoning processes. In this paper, we explore how self-play can enhance the rationality of models in the reasoning process without supervision from humans or superior models. We design a Critic-Discernment Game(CDG) in which a prover first provides a solution to a given problem and is subsequently challenged by critiques of its solution. These critiques either aim to assist or mislead the prover. The objective of the prover is to maintain the correct answer when faced with misleading comments, while correcting errors in response to constructive feedback. Our experiments on tasks involving mathematical reasoning, stepwise error detection, self-correction, and long-chain reasoning demonstrate that CDG training can significantly improve the ability of well-aligned LLMs to comprehend their reasoning process.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22920
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
Wang, Pinzheng
Li, Juntao
Tang, Zecheng
Gui, Haijia
zhang, Min
Artificial Intelligence
Large language models (LLMs) have demonstrated considerable reasoning abilities in various tasks such as mathematics and coding. However, recent studies indicate that even the best models lack true comprehension of their reasoning processes. In this paper, we explore how self-play can enhance the rationality of models in the reasoning process without supervision from humans or superior models. We design a Critic-Discernment Game(CDG) in which a prover first provides a solution to a given problem and is subsequently challenged by critiques of its solution. These critiques either aim to assist or mislead the prover. The objective of the prover is to maintain the correct answer when faced with misleading comments, while correcting errors in response to constructive feedback. Our experiments on tasks involving mathematical reasoning, stepwise error detection, self-correction, and long-chain reasoning demonstrate that CDG training can significantly improve the ability of well-aligned LLMs to comprehend their reasoning process.
title Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
topic Artificial Intelligence
url https://arxiv.org/abs/2506.22920