SSR: Socratic Self-Refine for Large Language Model Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Haizhou, Liu, Ye, Pang, Bo, Liu, Zeyu Leo, Wang, Hao, Savarese, Silvio, Xiong, Caiming, Zhou, Yingbo, Yavuz, Semih
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917079061364736
author Shi, Haizhou
Liu, Ye
Pang, Bo
Liu, Zeyu Leo
Wang, Hao
Savarese, Silvio
Xiong, Caiming
Zhou, Yingbo
Yavuz, Semih
author_facet Shi, Haizhou
Liu, Ye
Pang, Bo
Liu, Zeyu Leo
Wang, Hao
Savarese, Silvio
Xiong, Caiming
Zhou, Yingbo
Yavuz, Semih
contents Large Language Models (LLMs) have demonstrated remarkable reasoning abilities, yet existing test-time frameworks often rely on coarse self-verification and self-correction, limiting their effectiveness on complex tasks. In this paper, we propose Socratic Self-Refine (SSR), a novel framework for fine-grained evaluation and precise refinement of LLM reasoning. Our proposed SSR decomposes model responses into verifiable (sub-question, sub-answer) pairs, enabling step-level confidence estimation through controlled re-solving and self-consistency checks. By pinpointing unreliable steps and iteratively refining them, SSR produces more accurate and interpretable reasoning chains. Empirical results across five reasoning benchmarks and three LLMs show that SSR consistently outperforms state-of-the-art iterative self-refinement baselines. Beyond performance gains, SSR provides a principled black-box approach for evaluating and understanding the internal reasoning processes of LLMs. Code is available at https://github.com/SalesforceAIResearch/socratic-self-refine-reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SSR: Socratic Self-Refine for Large Language Model Reasoning
Shi, Haizhou
Liu, Ye
Pang, Bo
Liu, Zeyu Leo
Wang, Hao
Savarese, Silvio
Xiong, Caiming
Zhou, Yingbo
Yavuz, Semih
Computation and Language
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) have demonstrated remarkable reasoning abilities, yet existing test-time frameworks often rely on coarse self-verification and self-correction, limiting their effectiveness on complex tasks. In this paper, we propose Socratic Self-Refine (SSR), a novel framework for fine-grained evaluation and precise refinement of LLM reasoning. Our proposed SSR decomposes model responses into verifiable (sub-question, sub-answer) pairs, enabling step-level confidence estimation through controlled re-solving and self-consistency checks. By pinpointing unreliable steps and iteratively refining them, SSR produces more accurate and interpretable reasoning chains. Empirical results across five reasoning benchmarks and three LLMs show that SSR consistently outperforms state-of-the-art iterative self-refinement baselines. Beyond performance gains, SSR provides a principled black-box approach for evaluating and understanding the internal reasoning processes of LLMs. Code is available at https://github.com/SalesforceAIResearch/socratic-self-refine-reasoning.
title SSR: Socratic Self-Refine for Large Language Model Reasoning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.10621