PEARL: Plan Exploration and Adaptive Reinforcement Learning for Multihop Tool Use

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Qihao, Lu, Mingzhe, Wu, Jiayue, Hu, Yue, Liu, Yanbing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912854659039232
author Wang, Qihao
Lu, Mingzhe
Wu, Jiayue
Hu, Yue
Liu, Yanbing
author_facet Wang, Qihao
Lu, Mingzhe
Wu, Jiayue
Hu, Yue
Liu, Yanbing
contents Large Language Models show great potential with external tools, but face significant challenges in complex, multi-turn tool invocation. They often exhibit weak planning, tool hallucination, erroneous parameter generation, and struggle with robust interaction. To tackle these issues, we present PEARL, a novel framework to enhance LLM planning and execution for sophisticated tool use. PEARL adopts a two-stage approach: an offline phase where the agent explores tools to learn valid usage patterns and failure conditions, and an online reinforcement learning phase. In the online phase, a dedicated Planner is trained via group Relative Policy Optimization (GRPO) with a carefully designed reward function that provides distinct signals for planning quality. Experiments on the ToolHop and T-Eval benchmarks show PEARL significantly outperforms existing methods, achieving a new state-of-the-art success rate of \textbf{56.5\%} on ToolHop while maintaining a low invocation error rate. Our work marks a key advance in addressing the complex planning challenges of tool use, contributing to the development of more robust and reliable LLM-based agents.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20439
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PEARL: Plan Exploration and Adaptive Reinforcement Learning for Multihop Tool Use
Wang, Qihao
Lu, Mingzhe
Wu, Jiayue
Hu, Yue
Liu, Yanbing
Computation and Language
Large Language Models show great potential with external tools, but face significant challenges in complex, multi-turn tool invocation. They often exhibit weak planning, tool hallucination, erroneous parameter generation, and struggle with robust interaction. To tackle these issues, we present PEARL, a novel framework to enhance LLM planning and execution for sophisticated tool use. PEARL adopts a two-stage approach: an offline phase where the agent explores tools to learn valid usage patterns and failure conditions, and an online reinforcement learning phase. In the online phase, a dedicated Planner is trained via group Relative Policy Optimization (GRPO) with a carefully designed reward function that provides distinct signals for planning quality. Experiments on the ToolHop and T-Eval benchmarks show PEARL significantly outperforms existing methods, achieving a new state-of-the-art success rate of \textbf{56.5\%} on ToolHop while maintaining a low invocation error rate. Our work marks a key advance in addressing the complex planning challenges of tool use, contributing to the development of more robust and reliable LLM-based agents.
title PEARL: Plan Exploration and Adaptive Reinforcement Learning for Multihop Tool Use
topic Computation and Language
url https://arxiv.org/abs/2601.20439