Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909230410235904 |
|---|---|
| author | Komma, Abishek Chandrasekarasastry, Nagesh Panyam Leffel, Timothy Goyal, Anuj Metallinou, Angeliki Matsoukas, Spyros Galstyan, Aram |
| author_facet | Komma, Abishek Chandrasekarasastry, Nagesh Panyam Leffel, Timothy Goyal, Anuj Metallinou, Angeliki Matsoukas, Spyros Galstyan, Aram |
| contents | Measurement of interaction quality is a critical task for the improvement of spoken dialog systems. Existing approaches to dialog quality estimation either focus on evaluating the quality of individual turns, or collect dialog-level quality measurements from end users immediately following an interaction. In contrast to these approaches, we introduce a new dialog-level annotation workflow called Dialog Quality Annotation (DQA). DQA expert annotators evaluate the quality of dialogs as a whole, and also label dialogs for attributes such as goal completion and user sentiment. In this contribution, we show that: (i) while dialog quality cannot be completely decomposed into dialog-level attributes, there is a strong relationship between some objective dialog attributes and judgments of dialog quality; (ii) for the task of dialog-level quality estimation, a supervised model trained on dialog-level annotations outperforms methods based purely on aggregating turn-level features; and (iii) the proposed evaluation model shows better domain generalization ability compared to the baselines. On the basis of these results, we argue that having high-quality human-annotated data is an important component of evaluating interaction quality for large industrial-scale voice assistant platforms. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2306_03984 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs Komma, Abishek Chandrasekarasastry, Nagesh Panyam Leffel, Timothy Goyal, Anuj Metallinou, Angeliki Matsoukas, Spyros Galstyan, Aram Computation and Language Machine Learning Measurement of interaction quality is a critical task for the improvement of spoken dialog systems. Existing approaches to dialog quality estimation either focus on evaluating the quality of individual turns, or collect dialog-level quality measurements from end users immediately following an interaction. In contrast to these approaches, we introduce a new dialog-level annotation workflow called Dialog Quality Annotation (DQA). DQA expert annotators evaluate the quality of dialogs as a whole, and also label dialogs for attributes such as goal completion and user sentiment. In this contribution, we show that: (i) while dialog quality cannot be completely decomposed into dialog-level attributes, there is a strong relationship between some objective dialog attributes and judgments of dialog quality; (ii) for the task of dialog-level quality estimation, a supervised model trained on dialog-level annotations outperforms methods based purely on aggregating turn-level features; and (iii) the proposed evaluation model shows better domain generalization ability compared to the baselines. On the basis of these results, we argue that having high-quality human-annotated data is an important component of evaluating interaction quality for large industrial-scale voice assistant platforms. |
| title | Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2306.03984 |