Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhai, Yuanzhao, Yang, Tingkai, Xu, Kele, Dawei, Feng, Yang, Cheng, Ding, Bo, Wang, Huaimin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917775586361344
author Zhai, Yuanzhao
Yang, Tingkai
Xu, Kele
Dawei, Feng
Yang, Cheng
Ding, Bo
Wang, Huaimin
author_facet Zhai, Yuanzhao
Yang, Tingkai
Xu, Kele
Dawei, Feng
Yang, Cheng
Ding, Bo
Wang, Huaimin
contents Agents significantly enhance the capabilities of standalone Large Language Models (LLMs) by perceiving environments, making decisions, and executing actions. However, LLM agents still face challenges in tasks that require multiple decision-making steps. Estimating the value of actions in specific tasks is difficult when intermediate actions are neither appropriately rewarded nor penalized. In this paper, we propose leveraging a task-relevant Q-value model to guide action selection. Specifically, we first collect decision-making trajectories annotated with step-level Q values via Monte Carlo Tree Search (MCTS) and construct preference data. We then use another LLM to fit these preferences through step-level Direct Policy Optimization (DPO), which serves as the Q-value model. During inference, at each decision-making step, LLM agents select the action with the highest Q value before interacting with the environment. We apply our method to various open-source and API-based LLM agents, demonstrating that Q-value models significantly improve their performance. Notably, the performance of the agent built with Phi-3-mini-4k-instruct improved by 103% on WebShop and 75% on HotPotQA when enhanced with Q-value models, even surpassing GPT-4o-mini. Additionally, Q-value models offer several advantages, such as generalization to different LLM agents and seamless integration with existing prompting strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09345
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
Zhai, Yuanzhao
Yang, Tingkai
Xu, Kele
Dawei, Feng
Yang, Cheng
Ding, Bo
Wang, Huaimin
Artificial Intelligence
Agents significantly enhance the capabilities of standalone Large Language Models (LLMs) by perceiving environments, making decisions, and executing actions. However, LLM agents still face challenges in tasks that require multiple decision-making steps. Estimating the value of actions in specific tasks is difficult when intermediate actions are neither appropriately rewarded nor penalized. In this paper, we propose leveraging a task-relevant Q-value model to guide action selection. Specifically, we first collect decision-making trajectories annotated with step-level Q values via Monte Carlo Tree Search (MCTS) and construct preference data. We then use another LLM to fit these preferences through step-level Direct Policy Optimization (DPO), which serves as the Q-value model. During inference, at each decision-making step, LLM agents select the action with the highest Q value before interacting with the environment. We apply our method to various open-source and API-based LLM agents, demonstrating that Q-value models significantly improve their performance. Notably, the performance of the agent built with Phi-3-mini-4k-instruct improved by 103% on WebShop and 75% on HotPotQA when enhanced with Q-value models, even surpassing GPT-4o-mini. Additionally, Q-value models offer several advantages, such as generalization to different LLM agents and seamless integration with existing prompting strategies.
title Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
topic Artificial Intelligence
url https://arxiv.org/abs/2409.09345