Multi-Document Grounded Multi-Turn Synthetic Dialog Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Young-Suk, Gunasekara, Chulaka, Contractor, Danish, Astudillo, Ramón Fernandez, Florian, Radu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916399067168768
author Lee, Young-Suk
Gunasekara, Chulaka
Contractor, Danish
Astudillo, Ramón Fernandez
Florian, Radu
author_facet Lee, Young-Suk
Gunasekara, Chulaka
Contractor, Danish
Astudillo, Ramón Fernandez
Florian, Radu
contents We introduce a technique for multi-document grounded multi-turn synthetic dialog generation that incorporates three main ideas. First, we control the overall dialog flow using taxonomy-driven user queries that are generated with Chain-of-Thought (CoT) prompting. Second, we support the generation of multi-document grounded dialogs by mimicking real-world use of retrievers to update the grounding documents after every user-turn in the dialog. Third, we apply LLM-as-a-Judge to filter out queries with incorrect answers. Human evaluation of the synthetic dialog data suggests that the data is diverse, coherent, and includes mostly correct answers. Both human and automatic evaluations of answerable queries indicate that models fine-tuned on synthetic dialogs consistently out-perform those fine-tuned on existing human generated training data across four publicly available multi-turn document grounded benchmark test sets.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11500
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
Lee, Young-Suk
Gunasekara, Chulaka
Contractor, Danish
Astudillo, Ramón Fernandez
Florian, Radu
Computation and Language
Artificial Intelligence
We introduce a technique for multi-document grounded multi-turn synthetic dialog generation that incorporates three main ideas. First, we control the overall dialog flow using taxonomy-driven user queries that are generated with Chain-of-Thought (CoT) prompting. Second, we support the generation of multi-document grounded dialogs by mimicking real-world use of retrievers to update the grounding documents after every user-turn in the dialog. Third, we apply LLM-as-a-Judge to filter out queries with incorrect answers. Human evaluation of the synthetic dialog data suggests that the data is diverse, coherent, and includes mostly correct answers. Both human and automatic evaluations of answerable queries indicate that models fine-tuned on synthetic dialogs consistently out-perform those fine-tuned on existing human generated training data across four publicly available multi-turn document grounded benchmark test sets.
title Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.11500