Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kazi, Taaha, Lyu, Ruiliang, Zhou, Sizhe, Hakkani-Tur, Dilek, Tur, Gokhan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916482190934016
author Kazi, Taaha
Lyu, Ruiliang
Zhou, Sizhe
Hakkani-Tur, Dilek
Tur, Gokhan
author_facet Kazi, Taaha
Lyu, Ruiliang
Zhou, Sizhe
Hakkani-Tur, Dilek
Tur, Gokhan
contents Traditionally, offline datasets have been used to evaluate task-oriented dialogue (TOD) models. These datasets lack context awareness, making them suboptimal benchmarks for conversational systems. In contrast, user-agents, which are context-aware, can simulate the variability and unpredictability of human conversations, making them better alternatives as evaluators. Prior research has utilized large language models (LLMs) to develop user-agents. Our work builds upon this by using LLMs to create user-agents for the evaluation of TOD systems. This involves prompting an LLM, using in-context examples as guidance, and tracking the user-goal state. Our evaluation of diversity and task completion metrics for the user-agents shows improved performance with the use of better prompts. Additionally, we propose methodologies for the automatic evaluation of TOD models within this dynamic framework.
format Preprint
id arxiv_https___arxiv_org_abs_2411_09972
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
Kazi, Taaha
Lyu, Ruiliang
Zhou, Sizhe
Hakkani-Tur, Dilek
Tur, Gokhan
Computation and Language
Artificial Intelligence
Traditionally, offline datasets have been used to evaluate task-oriented dialogue (TOD) models. These datasets lack context awareness, making them suboptimal benchmarks for conversational systems. In contrast, user-agents, which are context-aware, can simulate the variability and unpredictability of human conversations, making them better alternatives as evaluators. Prior research has utilized large language models (LLMs) to develop user-agents. Our work builds upon this by using LLMs to create user-agents for the evaluation of TOD systems. This involves prompting an LLM, using in-context examples as guidance, and tracking the user-goal state. Our evaluation of diversity and task completion metrics for the user-agents shows improved performance with the use of better prompts. Additionally, we propose methodologies for the automatic evaluation of TOD models within this dynamic framework.
title Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.09972