Reasoning LLMs are Wandering Solution Explorers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Jiahao, Xu, Ziwei, Kankanhalli, Mohan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918034976800768
author Lu, Jiahao
Xu, Ziwei
Kankanhalli, Mohan
author_facet Lu, Jiahao
Xu, Ziwei
Kankanhalli, Mohan
contents Large Language Models (LLMs) have demonstrated impressive reasoning abilities through test-time computation (TTC) techniques such as chain-of-thought prompting and tree-based reasoning. However, we argue that current reasoning LLMs (RLLMs) lack the ability to systematically explore the solution space. This paper formalizes what constitutes systematic problem solving and identifies common failure modes that reveal reasoning LLMs to be wanderers rather than systematic explorers. Through qualitative and quantitative analysis across multiple state-of-the-art LLMs, we uncover persistent issues: invalid reasoning steps, redundant explorations, hallucinated or unfaithful conclusions, and so on. Our findings suggest that current models' performance can appear to be competent on simple tasks yet degrade sharply as complexity increases. Based on the findings, we advocate for new metrics and tools that evaluate not just final outputs but the structure of the reasoning process itself.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20296
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning LLMs are Wandering Solution Explorers
Lu, Jiahao
Xu, Ziwei
Kankanhalli, Mohan
Computation and Language
Artificial Intelligence
Machine Learning
Multimedia
Large Language Models (LLMs) have demonstrated impressive reasoning abilities through test-time computation (TTC) techniques such as chain-of-thought prompting and tree-based reasoning. However, we argue that current reasoning LLMs (RLLMs) lack the ability to systematically explore the solution space. This paper formalizes what constitutes systematic problem solving and identifies common failure modes that reveal reasoning LLMs to be wanderers rather than systematic explorers. Through qualitative and quantitative analysis across multiple state-of-the-art LLMs, we uncover persistent issues: invalid reasoning steps, redundant explorations, hallucinated or unfaithful conclusions, and so on. Our findings suggest that current models' performance can appear to be competent on simple tasks yet degrade sharply as complexity increases. Based on the findings, we advocate for new metrics and tools that evaluate not just final outputs but the structure of the reasoning process itself.
title Reasoning LLMs are Wandering Solution Explorers
topic Computation and Language
Artificial Intelligence
Machine Learning
Multimedia
url https://arxiv.org/abs/2505.20296