From Implicit Exploration to Structured Reasoning: Leveraging Guideline and Refinement for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jiaxiang, Wang, Zhuo, Zou, Mingxi, Li, Zhucong, Zhou, Zhijian, Wang, Song, Xu, Zenglin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908524522504192
author Chen, Jiaxiang
Wang, Zhuo
Zou, Mingxi
Li, Zhucong
Zhou, Zhijian
Wang, Song
Xu, Zenglin
author_facet Chen, Jiaxiang
Wang, Zhuo
Zou, Mingxi
Li, Zhucong
Zhou, Zhijian
Wang, Song
Xu, Zenglin
contents Large language models (LLMs) have advanced general-purpose reasoning, showing strong performance across diverse tasks. However, existing methods often rely on implicit exploration, where the model follows stochastic and unguided reasoning paths-like walking without a map. This leads to unstable reasoning paths, lack of error correction, and limited learning from past experience. To address these issues, we propose a framework that shifts from implicit exploration to structured reasoning through guideline and refinement. First, we extract structured reasoning patterns from successful trajectories and reflective signals from failures. During inference, the model follows these guidelines step-by-step, with refinement applied after each step to correct errors and stabilize the reasoning process. Experiments on BBH and four additional benchmarks (GSM8K, MATH-500, MBPP, HumanEval) show that our method consistently outperforms strong baselines across diverse reasoning tasks. Structured reasoning with stepwise execution and refinement improves stability and generalization, while guidelines transfer well across domains and flexibly support cross-model collaboration, matching or surpassing supervised fine-tuning in effectiveness and scalability.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Implicit Exploration to Structured Reasoning: Leveraging Guideline and Refinement for LLMs
Chen, Jiaxiang
Wang, Zhuo
Zou, Mingxi
Li, Zhucong
Zhou, Zhijian
Wang, Song
Xu, Zenglin
Artificial Intelligence
Machine Learning
Large language models (LLMs) have advanced general-purpose reasoning, showing strong performance across diverse tasks. However, existing methods often rely on implicit exploration, where the model follows stochastic and unguided reasoning paths-like walking without a map. This leads to unstable reasoning paths, lack of error correction, and limited learning from past experience. To address these issues, we propose a framework that shifts from implicit exploration to structured reasoning through guideline and refinement. First, we extract structured reasoning patterns from successful trajectories and reflective signals from failures. During inference, the model follows these guidelines step-by-step, with refinement applied after each step to correct errors and stabilize the reasoning process. Experiments on BBH and four additional benchmarks (GSM8K, MATH-500, MBPP, HumanEval) show that our method consistently outperforms strong baselines across diverse reasoning tasks. Structured reasoning with stepwise execution and refinement improves stability and generalization, while guidelines transfer well across domains and flexibly support cross-model collaboration, matching or surpassing supervised fine-tuning in effectiveness and scalability.
title From Implicit Exploration to Structured Reasoning: Leveraging Guideline and Refinement for LLMs
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.06284