Learning from Relevant Subgoals in Successful Dialogs using Iterative Training for Task-oriented Dialog Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kaiser, Magdalena, Ernst, Patrick, Szarvas, György
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909403848900608
author Kaiser, Magdalena
Ernst, Patrick
Szarvas, György
author_facet Kaiser, Magdalena
Ernst, Patrick
Szarvas, György
contents Task-oriented Dialog (ToD) systems have to solve multiple subgoals to accomplish user goals, whereas feedback is often obtained only at the end of the dialog. In this work, we propose SUIT (SUbgoal-aware ITerative Training), an iterative training approach for improving ToD systems. We sample dialogs from the model we aim to improve and determine subgoals that contribute to dialog success using distant supervision to obtain high quality training samples. We show how this data improves supervised fine-tuning or, alternatively, preference learning results. SUIT is able to iteratively generate more data instead of relying on fixed static sets. SUIT reaches new state-of-the-art performance on a popular ToD benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16305
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning from Relevant Subgoals in Successful Dialogs using Iterative Training for Task-oriented Dialog Systems
Kaiser, Magdalena
Ernst, Patrick
Szarvas, György
Computation and Language
Artificial Intelligence
Machine Learning
Task-oriented Dialog (ToD) systems have to solve multiple subgoals to accomplish user goals, whereas feedback is often obtained only at the end of the dialog. In this work, we propose SUIT (SUbgoal-aware ITerative Training), an iterative training approach for improving ToD systems. We sample dialogs from the model we aim to improve and determine subgoals that contribute to dialog success using distant supervision to obtain high quality training samples. We show how this data improves supervised fine-tuning or, alternatively, preference learning results. SUIT is able to iteratively generate more data instead of relying on fixed static sets. SUIT reaches new state-of-the-art performance on a popular ToD benchmark.
title Learning from Relevant Subgoals in Successful Dialogs using Iterative Training for Task-oriented Dialog Systems
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.16305