VoiceBench: Benchmarking LLM-Based Voice Assistants

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yiming, Yue, Xianghu, Zhang, Chen, Gao, Xiaoxue, Tan, Robby T., Li, Haizhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916518712836096
author Chen, Yiming
Yue, Xianghu
Zhang, Chen
Gao, Xiaoxue
Tan, Robby T.
Li, Haizhou
author_facet Chen, Yiming
Yue, Xianghu
Zhang, Chen
Gao, Xiaoxue
Tan, Robby T.
Li, Haizhou
contents Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience compared to traditional text-based interactions. However, the absence of benchmarks designed to evaluate these speech interaction capabilities has hindered progress of LLM-based voice assistants development. Current evaluations focus primarily on automatic speech recognition (ASR) or general knowledge evaluation with clean speeches, neglecting the more intricate, real-world scenarios that involve diverse speaker characteristics, environmental and content factors. To address this, we introduce VoiceBench, the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants. VoiceBench also includes both real and synthetic spoken instructions that incorporate the above three key real-world variations. Extensive experiments reveal the limitations of current LLM-based voice assistant models and offer valuable insights for future research and development in this field.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17196
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VoiceBench: Benchmarking LLM-Based Voice Assistants
Chen, Yiming
Yue, Xianghu
Zhang, Chen
Gao, Xiaoxue
Tan, Robby T.
Li, Haizhou
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience compared to traditional text-based interactions. However, the absence of benchmarks designed to evaluate these speech interaction capabilities has hindered progress of LLM-based voice assistants development. Current evaluations focus primarily on automatic speech recognition (ASR) or general knowledge evaluation with clean speeches, neglecting the more intricate, real-world scenarios that involve diverse speaker characteristics, environmental and content factors. To address this, we introduce VoiceBench, the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants. VoiceBench also includes both real and synthetic spoken instructions that incorporate the above three key real-world variations. Extensive experiments reveal the limitations of current LLM-based voice assistant models and offer valuable insights for future research and development in this field.
title VoiceBench: Benchmarking LLM-Based Voice Assistants
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.17196