LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Stacey, Joe, Cheng, Jianpeng, Torr, John, Guigue, Tristan, Driesen, Joris, Coca, Alexandru, Gaynor, Mark, Johannsen, Anders
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910432929775616
author Stacey, Joe
Cheng, Jianpeng
Torr, John
Guigue, Tristan
Driesen, Joris
Coca, Alexandru
Gaynor, Mark
Johannsen, Anders
author_facet Stacey, Joe
Cheng, Jianpeng
Torr, John
Guigue, Tristan
Driesen, Joris
Coca, Alexandru
Gaynor, Mark
Johannsen, Anders
contents Spurred by recent advances in Large Language Models (LLMs), virtual assistants are poised to take a leap forward in terms of their dialogue capabilities. Yet a major bottleneck to achieving genuinely transformative task-oriented dialogue capabilities remains the scarcity of high quality data. Existing datasets, while impressive in scale, have limited domain coverage and contain few genuinely challenging conversational phenomena; those which are present are typically unlabelled, making it difficult to assess the strengths and weaknesses of models without time-consuming and costly human evaluation. Moreover, creating high quality dialogue data has until now required considerable human input, limiting both the scale of these datasets and the ability to rapidly bootstrap data for a new target domain. We aim to overcome these issues with LUCID, a modularised and highly automated LLM-driven data generation system that produces realistic, diverse and challenging dialogues. We use LUCID to generate a seed dataset of 4,277 conversations across 100 intents to demonstrate its capabilities, with a human review finding consistently high quality labels in the generated data.
format Preprint
id arxiv_https___arxiv_org_abs_2403_00462
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues
Stacey, Joe
Cheng, Jianpeng
Torr, John
Guigue, Tristan
Driesen, Joris
Coca, Alexandru
Gaynor, Mark
Johannsen, Anders
Computation and Language
I.2.7
Spurred by recent advances in Large Language Models (LLMs), virtual assistants are poised to take a leap forward in terms of their dialogue capabilities. Yet a major bottleneck to achieving genuinely transformative task-oriented dialogue capabilities remains the scarcity of high quality data. Existing datasets, while impressive in scale, have limited domain coverage and contain few genuinely challenging conversational phenomena; those which are present are typically unlabelled, making it difficult to assess the strengths and weaknesses of models without time-consuming and costly human evaluation. Moreover, creating high quality dialogue data has until now required considerable human input, limiting both the scale of these datasets and the ability to rapidly bootstrap data for a new target domain. We aim to overcome these issues with LUCID, a modularised and highly automated LLM-driven data generation system that produces realistic, diverse and challenging dialogues. We use LUCID to generate a seed dataset of 4,277 conversations across 100 intents to demonstrate its capabilities, with a human review finding consistently high quality labels in the generated data.
title LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2403.00462