SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Si, Shuzheng, Ma, Wentao, Gao, Haoyu, Wu, Yuchuan, Lin, Ting-En, Dai, Yinpei, Li, Hangyu, Yan, Rui, Huang, Fei, Li, Yongbin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911019809374208
author Si, Shuzheng
Ma, Wentao
Gao, Haoyu
Wu, Yuchuan
Lin, Ting-En
Dai, Yinpei
Li, Hangyu
Yan, Rui
Huang, Fei
Li, Yongbin
author_facet Si, Shuzheng
Ma, Wentao
Gao, Haoyu
Wu, Yuchuan
Lin, Ting-En
Dai, Yinpei
Li, Hangyu
Yan, Rui
Huang, Fei
Li, Yongbin
contents Task-oriented dialogue (TOD) models have made significant progress in recent years. However, previous studies primarily focus on datasets written by annotators, which has resulted in a gap between academic research and real-world spoken conversation scenarios. While several small-scale spoken TOD datasets are proposed to address robustness issues such as ASR errors, they ignore the unique challenges in spoken conversation. To tackle the limitations, we introduce SpokenWOZ, a large-scale speech-text dataset for spoken TOD, containing 8 domains, 203k turns, 5.7k dialogues and 249 hours of audios from human-to-human spoken conversations. SpokenWOZ further incorporates common spoken characteristics such as word-by-word processing and reasoning in spoken language. Based on these characteristics, we present cross-turn slot and reasoning slot detection as new challenges. We conduct experiments on various baselines, including text-modal models, newly proposed dual-modal models, and LLMs, e.g., ChatGPT. The results show that the current models still have substantial room for improvement in spoken conversation, where the most advanced dialogue state tracker only achieves 25.65% in joint goal accuracy and the SOTA end-to-end model only correctly completes the user request in 52.1% of dialogues. The dataset, code, and leaderboard are available: https://spokenwoz.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2305_13040
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
Si, Shuzheng
Ma, Wentao
Gao, Haoyu
Wu, Yuchuan
Lin, Ting-En
Dai, Yinpei
Li, Hangyu
Yan, Rui
Huang, Fei
Li, Yongbin
Computation and Language
Artificial Intelligence
Task-oriented dialogue (TOD) models have made significant progress in recent years. However, previous studies primarily focus on datasets written by annotators, which has resulted in a gap between academic research and real-world spoken conversation scenarios. While several small-scale spoken TOD datasets are proposed to address robustness issues such as ASR errors, they ignore the unique challenges in spoken conversation. To tackle the limitations, we introduce SpokenWOZ, a large-scale speech-text dataset for spoken TOD, containing 8 domains, 203k turns, 5.7k dialogues and 249 hours of audios from human-to-human spoken conversations. SpokenWOZ further incorporates common spoken characteristics such as word-by-word processing and reasoning in spoken language. Based on these characteristics, we present cross-turn slot and reasoning slot detection as new challenges. We conduct experiments on various baselines, including text-modal models, newly proposed dual-modal models, and LLMs, e.g., ChatGPT. The results show that the current models still have substantial room for improvement in spoken conversation, where the most advanced dialogue state tracker only achieves 25.65% in joint goal accuracy and the SOTA end-to-end model only correctly completes the user request in 52.1% of dialogues. The dataset, code, and leaderboard are available: https://spokenwoz.github.io/.
title SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2305.13040