ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhigen, Peng, Jianxiang, Wang, Yanmeng, Cao, Yong, Shen, Tianhao, Zhang, Minghui, Su, Linxi, Wu, Shang, Wu, Yihang, Wang, Yuqian, Wang, Ye, Hu, Wei, Li, Jianfeng, Wang, Shaojun, Xiao, Jing, Xiong, Deyi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929725423747072
author Li, Zhigen
Peng, Jianxiang
Wang, Yanmeng
Cao, Yong
Shen, Tianhao
Zhang, Minghui
Su, Linxi
Wu, Shang
Wu, Yihang
Wang, Yuqian
Wang, Ye
Hu, Wei
Li, Jianfeng
Wang, Shaojun
Xiao, Jing
Xiong, Deyi
author_facet Li, Zhigen
Peng, Jianxiang
Wang, Yanmeng
Cao, Yong
Shen, Tianhao
Zhang, Minghui
Su, Linxi
Wu, Shang
Wu, Yihang
Wang, Yuqian
Wang, Ye
Hu, Wei
Li, Jianfeng
Wang, Shaojun
Xiao, Jing
Xiong, Deyi
contents Dialogue agents powered by Large Language Models (LLMs) show superior performance in various tasks. Despite the better user understanding and human-like responses, their lack of controllability remains a key challenge, often leading to unfocused conversations or task failure. To address this, we introduce Standard Operating Procedure (SOP) to regulate dialogue flow. Specifically, we propose ChatSOP, a novel SOP-guided Monte Carlo Tree Search (MCTS) planning framework designed to enhance the controllability of LLM-driven dialogue agents. To enable this, we curate a dataset comprising SOP-annotated multi-scenario dialogues, generated using a semi-automated role-playing system with GPT-4o and validated through strict manual quality control. Additionally, we propose a novel method that integrates Chain of Thought reasoning with supervised fine-tuning for SOP prediction and utilizes SOP-guided Monte Carlo Tree Search for optimal action planning during dialogues. Experimental results demonstrate the effectiveness of our method, such as achieving a 27.95% improvement in action accuracy compared to baseline models based on GPT-3.5 and also showing notable gains for open-source models. Dataset and codes are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03884
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
Li, Zhigen
Peng, Jianxiang
Wang, Yanmeng
Cao, Yong
Shen, Tianhao
Zhang, Minghui
Su, Linxi
Wu, Shang
Wu, Yihang
Wang, Yuqian
Wang, Ye
Hu, Wei
Li, Jianfeng
Wang, Shaojun
Xiao, Jing
Xiong, Deyi
Computation and Language
Artificial Intelligence
Dialogue agents powered by Large Language Models (LLMs) show superior performance in various tasks. Despite the better user understanding and human-like responses, their lack of controllability remains a key challenge, often leading to unfocused conversations or task failure. To address this, we introduce Standard Operating Procedure (SOP) to regulate dialogue flow. Specifically, we propose ChatSOP, a novel SOP-guided Monte Carlo Tree Search (MCTS) planning framework designed to enhance the controllability of LLM-driven dialogue agents. To enable this, we curate a dataset comprising SOP-annotated multi-scenario dialogues, generated using a semi-automated role-playing system with GPT-4o and validated through strict manual quality control. Additionally, we propose a novel method that integrates Chain of Thought reasoning with supervised fine-tuning for SOP prediction and utilizes SOP-guided Monte Carlo Tree Search for optimal action planning during dialogues. Experimental results demonstrate the effectiveness of our method, such as achieving a 27.95% improvement in action accuracy compared to baseline models based on GPT-3.5 and also showing notable gains for open-source models. Dataset and codes are publicly available.
title ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.03884