Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zehong, Wu, Fang, Wang, Hongru, Tang, Xiangru, Li, Bolian, Yin, Zhenfei, Ma, Yijun, Li, Yiyang, Sun, Weixiang, Chen, Xiusi, Ye, Yanfang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915761966022656
author Wang, Zehong
Wu, Fang
Wang, Hongru
Tang, Xiangru
Li, Bolian
Yin, Zhenfei
Ma, Yijun
Li, Yiyang
Sun, Weixiang
Chen, Xiusi
Ye, Yanfang
author_facet Wang, Zehong
Wu, Fang
Wang, Hongru
Tang, Xiangru
Li, Bolian
Yin, Zhenfei
Ma, Yijun
Li, Yiyang
Sun, Weixiang
Chen, Xiusi
Ye, Yanfang
contents Large language model (LLM)-based agents exhibit strong step-by-step reasoning capabilities over short horizons, yet often fail to sustain coherent behavior over long planning horizons. We argue that this failure reflects a fundamental mismatch: step-wise reasoning induces a form of step-wise greedy policy that is adequate for short horizons but fails in long-horizon planning, where early actions must account for delayed consequences. From this planning-centric perspective, we study LLM-based agents in deterministic, fully structured environments with explicit state transitions and evaluation signals. Our analysis reveals a core failure mode of reasoning-based policies: locally optimal choices induced by step-wise scoring lead to early myopic commitments that are systematically amplified over time and difficult to recover from. We introduce FLARE (Future-aware Lookahead with Reward Estimation) as a minimal instantiation of future-aware planning to enforce explicit lookahead, value propagation, and limited commitment in a single model, allowing downstream outcomes to influence early decisions. Across multiple benchmarks, agent frameworks, and LLM backbones, FLARE consistently improves task performance and planning-level behavior, frequently allowing LLaMA-8B with FLARE to outperform GPT-4o with standard step-by-step reasoning. These results establish a clear distinction between reasoning and planning.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22311
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
Wang, Zehong
Wu, Fang
Wang, Hongru
Tang, Xiangru
Li, Bolian
Yin, Zhenfei
Ma, Yijun
Li, Yiyang
Sun, Weixiang
Chen, Xiusi
Ye, Yanfang
Artificial Intelligence
Computation and Language
Machine Learning
Large language model (LLM)-based agents exhibit strong step-by-step reasoning capabilities over short horizons, yet often fail to sustain coherent behavior over long planning horizons. We argue that this failure reflects a fundamental mismatch: step-wise reasoning induces a form of step-wise greedy policy that is adequate for short horizons but fails in long-horizon planning, where early actions must account for delayed consequences. From this planning-centric perspective, we study LLM-based agents in deterministic, fully structured environments with explicit state transitions and evaluation signals. Our analysis reveals a core failure mode of reasoning-based policies: locally optimal choices induced by step-wise scoring lead to early myopic commitments that are systematically amplified over time and difficult to recover from. We introduce FLARE (Future-aware Lookahead with Reward Estimation) as a minimal instantiation of future-aware planning to enforce explicit lookahead, value propagation, and limited commitment in a single model, allowing downstream outcomes to influence early decisions. Across multiple benchmarks, agent frameworks, and LLM backbones, FLARE consistently improves task performance and planning-level behavior, frequently allowing LLaMA-8B with FLARE to outperform GPT-4o with standard step-by-step reasoning. These results establish a clear distinction between reasoning and planning.
title Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.22311