Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Komma, Abishek, Chandrasekarasastry, Nagesh Panyam, Leffel, Timothy, Goyal, Anuj, Metallinou, Angeliki, Matsoukas, Spyros, Galstyan, Aram
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909230410235904
author Komma, Abishek
Chandrasekarasastry, Nagesh Panyam
Leffel, Timothy
Goyal, Anuj
Metallinou, Angeliki
Matsoukas, Spyros
Galstyan, Aram
author_facet Komma, Abishek
Chandrasekarasastry, Nagesh Panyam
Leffel, Timothy
Goyal, Anuj
Metallinou, Angeliki
Matsoukas, Spyros
Galstyan, Aram
contents Measurement of interaction quality is a critical task for the improvement of spoken dialog systems. Existing approaches to dialog quality estimation either focus on evaluating the quality of individual turns, or collect dialog-level quality measurements from end users immediately following an interaction. In contrast to these approaches, we introduce a new dialog-level annotation workflow called Dialog Quality Annotation (DQA). DQA expert annotators evaluate the quality of dialogs as a whole, and also label dialogs for attributes such as goal completion and user sentiment. In this contribution, we show that: (i) while dialog quality cannot be completely decomposed into dialog-level attributes, there is a strong relationship between some objective dialog attributes and judgments of dialog quality; (ii) for the task of dialog-level quality estimation, a supervised model trained on dialog-level annotations outperforms methods based purely on aggregating turn-level features; and (iii) the proposed evaluation model shows better domain generalization ability compared to the baselines. On the basis of these results, we argue that having high-quality human-annotated data is an important component of evaluating interaction quality for large industrial-scale voice assistant platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2306_03984
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs
Komma, Abishek
Chandrasekarasastry, Nagesh Panyam
Leffel, Timothy
Goyal, Anuj
Metallinou, Angeliki
Matsoukas, Spyros
Galstyan, Aram
Computation and Language
Machine Learning
Measurement of interaction quality is a critical task for the improvement of spoken dialog systems. Existing approaches to dialog quality estimation either focus on evaluating the quality of individual turns, or collect dialog-level quality measurements from end users immediately following an interaction. In contrast to these approaches, we introduce a new dialog-level annotation workflow called Dialog Quality Annotation (DQA). DQA expert annotators evaluate the quality of dialogs as a whole, and also label dialogs for attributes such as goal completion and user sentiment. In this contribution, we show that: (i) while dialog quality cannot be completely decomposed into dialog-level attributes, there is a strong relationship between some objective dialog attributes and judgments of dialog quality; (ii) for the task of dialog-level quality estimation, a supervised model trained on dialog-level annotations outperforms methods based purely on aggregating turn-level features; and (iii) the proposed evaluation model shows better domain generalization ability compared to the baselines. On the basis of these results, we argue that having high-quality human-annotated data is an important component of evaluating interaction quality for large industrial-scale voice assistant platforms.
title Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2306.03984