WavChat: A Survey of Spoken Dialogue Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ji, Shengpeng, Chen, Yifu, Fang, Minghui, Zuo, Jialong, Lu, Jingyu, Wang, Hanting, Jiang, Ziyue, Zhou, Long, Liu, Shujie, Cheng, Xize, Yang, Xiaoda, Wang, Zehan, Yang, Qian, Li, Jian, Jiang, Yidi, He, Jingzhen, Chu, Yunfei, Xu, Jin, Zhao, Zhou
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929605553684480
author Ji, Shengpeng
Chen, Yifu
Fang, Minghui
Zuo, Jialong
Lu, Jingyu
Wang, Hanting
Jiang, Ziyue
Zhou, Long
Liu, Shujie
Cheng, Xize
Yang, Xiaoda
Wang, Zehan
Yang, Qian
Li, Jian
Jiang, Yidi
He, Jingzhen
Chu, Yunfei
Xu, Jin
Zhao, Zhou
author_facet Ji, Shengpeng
Chen, Yifu
Fang, Minghui
Zuo, Jialong
Lu, Jingyu
Wang, Hanting
Jiang, Ziyue
Zhou, Long
Liu, Shujie
Cheng, Xize
Yang, Xiaoda
Wang, Zehan
Yang, Qian
Li, Jian
Jiang, Yidi
He, Jingzhen
Chu, Yunfei
Xu, Jin
Zhao, Zhou
contents Recent advancements in spoken dialogue models, exemplified by systems like GPT-4o, have captured significant attention in the speech domain. Compared to traditional three-tier cascaded spoken dialogue models that comprise speech recognition (ASR), large language models (LLMs), and text-to-speech (TTS), modern spoken dialogue models exhibit greater intelligence. These advanced spoken dialogue models not only comprehend audio, music, and other speech-related features, but also capture stylistic and timbral characteristics in speech. Moreover, they generate high-quality, multi-turn speech responses with low latency, enabling real-time interaction through simultaneous listening and speaking capability. Despite the progress in spoken dialogue systems, there is a lack of comprehensive surveys that systematically organize and analyze these systems and the underlying technologies. To address this, we have first compiled existing spoken dialogue systems in the chronological order and categorized them into the cascaded and end-to-end paradigms. We then provide an in-depth overview of the core technologies in spoken dialogue models, covering aspects such as speech representation, training paradigm, streaming, duplex, and interaction capabilities. Each section discusses the limitations of these technologies and outlines considerations for future research. Additionally, we present a thorough review of relevant datasets, evaluation metrics, and benchmarks from the perspectives of training and evaluating spoken dialogue systems. We hope this survey will contribute to advancing both academic research and industrial applications in the field of spoken dialogue systems. The related material is available at https://github.com/jishengpeng/WavChat.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13577
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WavChat: A Survey of Spoken Dialogue Models
Ji, Shengpeng
Chen, Yifu
Fang, Minghui
Zuo, Jialong
Lu, Jingyu
Wang, Hanting
Jiang, Ziyue
Zhou, Long
Liu, Shujie
Cheng, Xize
Yang, Xiaoda
Wang, Zehan
Yang, Qian
Li, Jian
Jiang, Yidi
He, Jingzhen
Chu, Yunfei
Xu, Jin
Zhao, Zhou
Audio and Speech Processing
Computation and Language
Machine Learning
Multimedia
Sound
Recent advancements in spoken dialogue models, exemplified by systems like GPT-4o, have captured significant attention in the speech domain. Compared to traditional three-tier cascaded spoken dialogue models that comprise speech recognition (ASR), large language models (LLMs), and text-to-speech (TTS), modern spoken dialogue models exhibit greater intelligence. These advanced spoken dialogue models not only comprehend audio, music, and other speech-related features, but also capture stylistic and timbral characteristics in speech. Moreover, they generate high-quality, multi-turn speech responses with low latency, enabling real-time interaction through simultaneous listening and speaking capability. Despite the progress in spoken dialogue systems, there is a lack of comprehensive surveys that systematically organize and analyze these systems and the underlying technologies. To address this, we have first compiled existing spoken dialogue systems in the chronological order and categorized them into the cascaded and end-to-end paradigms. We then provide an in-depth overview of the core technologies in spoken dialogue models, covering aspects such as speech representation, training paradigm, streaming, duplex, and interaction capabilities. Each section discusses the limitations of these technologies and outlines considerations for future research. Additionally, we present a thorough review of relevant datasets, evaluation metrics, and benchmarks from the perspectives of training and evaluating spoken dialogue systems. We hope this survey will contribute to advancing both academic research and industrial applications in the field of spoken dialogue systems. The related material is available at https://github.com/jishengpeng/WavChat.
title WavChat: A Survey of Spoken Dialogue Models
topic Audio and Speech Processing
Computation and Language
Machine Learning
Multimedia
Sound
url https://arxiv.org/abs/2411.13577