Convert Language Model into a Value-based Strategic Planner

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiaoyu, Zhao, Yue, Gu, Qingqing, Jiang, Zhonglin, Chen, Xiaokai, Chen, Yong, Ji, Luo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915465039708160
author Wang, Xiaoyu
Zhao, Yue
Gu, Qingqing
Jiang, Zhonglin
Chen, Xiaokai
Chen, Yong
Ji, Luo
author_facet Wang, Xiaoyu
Zhao, Yue
Gu, Qingqing
Jiang, Zhonglin
Chen, Xiaokai
Chen, Yong
Ji, Luo
contents Emotional support conversation (ESC) aims to alleviate the emotional distress of individuals through effective conversations. Although large language models (LLMs) have obtained remarkable progress on ESC, most of these studies might not define the diagram from the state model perspective, therefore providing a suboptimal solution for long-term satisfaction. To address such an issue, we leverage the Q-learning on LLMs, and propose a framework called straQ*. Our framework allows a plug-and-play LLM to bootstrap the planning during ESC, determine the optimal strategy based on long-term returns, and finally guide the LLM to response. Substantial experiments on ESC datasets suggest that straQ* outperforms many baselines, including direct inference, self-refine, chain of thought, finetuning, and finite state machines.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06987
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Convert Language Model into a Value-based Strategic Planner
Wang, Xiaoyu
Zhao, Yue
Gu, Qingqing
Jiang, Zhonglin
Chen, Xiaokai
Chen, Yong
Ji, Luo
Computation and Language
Artificial Intelligence
Emotional support conversation (ESC) aims to alleviate the emotional distress of individuals through effective conversations. Although large language models (LLMs) have obtained remarkable progress on ESC, most of these studies might not define the diagram from the state model perspective, therefore providing a suboptimal solution for long-term satisfaction. To address such an issue, we leverage the Q-learning on LLMs, and propose a framework called straQ*. Our framework allows a plug-and-play LLM to bootstrap the planning during ESC, determine the optimal strategy based on long-term returns, and finally guide the LLM to response. Substantial experiments on ESC datasets suggest that straQ* outperforms many baselines, including direct inference, self-refine, chain of thought, finetuning, and finite state machines.
title Convert Language Model into a Value-based Strategic Planner
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.06987