Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Siwei, Shen, Yifei, Sun, Haoran, Feng, Shi, Teng, Shang-Hua, Dong, Li, Hao, Yaru, Chen, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917307233599488
author Wang, Siwei
Shen, Yifei
Sun, Haoran
Feng, Shi
Teng, Shang-Hua
Dong, Li
Hao, Yaru
Chen, Wei
author_facet Wang, Siwei
Shen, Yifei
Sun, Haoran
Feng, Shi
Teng, Shang-Hua
Dong, Li
Hao, Yaru
Chen, Wei
contents Recent reinforcement learning (RL) methods have substantially enhanced the planning capabilities of Large Language Models (LLMs), yet the theoretical basis for their effectiveness remains elusive. In this work, we investigate RL's benefits and limitations through a tractable graph-based abstraction, focusing on policy gradient (PG) and Q-learning methods. Our theoretical analyses reveal that supervised fine-tuning (SFT) may introduce co-occurrence-based spurious solutions, whereas RL achieves correct planning primarily through exploration, underscoring exploration's role in enabling better generalization. However, we also show that PG suffers from diversity collapse, where output diversity decreases during training and persists even after perfect accuracy is attained. By contrast, Q-learning provides two key advantages: off-policy learning and diversity preservation at convergence. We further demonstrate that careful reward design is necessary to prevent Q-value bias in Q-learning. Finally, applying our framework to the real-world planning benchmark Blocksworld, we confirm that these behaviors manifest in practice.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22613
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
Wang, Siwei
Shen, Yifei
Sun, Haoran
Feng, Shi
Teng, Shang-Hua
Dong, Li
Hao, Yaru
Chen, Wei
Artificial Intelligence
Computation and Language
Machine Learning
Recent reinforcement learning (RL) methods have substantially enhanced the planning capabilities of Large Language Models (LLMs), yet the theoretical basis for their effectiveness remains elusive. In this work, we investigate RL's benefits and limitations through a tractable graph-based abstraction, focusing on policy gradient (PG) and Q-learning methods. Our theoretical analyses reveal that supervised fine-tuning (SFT) may introduce co-occurrence-based spurious solutions, whereas RL achieves correct planning primarily through exploration, underscoring exploration's role in enabling better generalization. However, we also show that PG suffers from diversity collapse, where output diversity decreases during training and persists even after perfect accuracy is attained. By contrast, Q-learning provides two key advantages: off-policy learning and diversity preservation at convergence. We further demonstrate that careful reward design is necessary to prevent Q-value bias in Q-learning. Finally, applying our framework to the real-world planning benchmark Blocksworld, we confirm that these behaviors manifest in practice.
title Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.22613