VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cui, Wenqian, Jiao, Xiaoqi, Meng, Ziqiao, King, Irwin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918035855507456
author Cui, Wenqian
Jiao, Xiaoqi
Meng, Ziqiao
King, Irwin
author_facet Cui, Wenqian
Jiao, Xiaoqi
Meng, Ziqiao
King, Irwin
contents With the rising need for speech-based interaction models, end-to-end Spoken Language Models (SLMs) have emerged as a promising solution. While these models require comprehensive world knowledge for meaningful and reliable human interactions, existing question-answering (QA) benchmarks fall short in evaluating SLMs' knowledge understanding due to their inability to support end-to-end speech evaluation and account for varied input audio conditions. To address these limitations, we present VoxEval, a novel SpeechQA benchmark that assesses SLMs' knowledge understanding through pure speech interactions. Our benchmark 1) uniquely maintains speech format for both inputs and outputs, 2) evaluates model robustness across diverse input audio conditions, and 3) pioneers the assessment of complex tasks like mathematical reasoning in spoken format. Systematic evaluation demonstrates that VoxEval presents significant challenges to current SLMs, revealing their sensitivity to varying audio conditions and highlighting the need to enhance reasoning capabilities in future development. We hope this benchmark could guide the advancement of more sophisticated and reliable SLMs. VoxEval dataset is available at: https://github.com/dreamtheater123/VoxEval
format Preprint
id arxiv_https___arxiv_org_abs_2501_04962
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
Cui, Wenqian
Jiao, Xiaoqi
Meng, Ziqiao
King, Irwin
Computation and Language
Sound
Audio and Speech Processing
With the rising need for speech-based interaction models, end-to-end Spoken Language Models (SLMs) have emerged as a promising solution. While these models require comprehensive world knowledge for meaningful and reliable human interactions, existing question-answering (QA) benchmarks fall short in evaluating SLMs' knowledge understanding due to their inability to support end-to-end speech evaluation and account for varied input audio conditions. To address these limitations, we present VoxEval, a novel SpeechQA benchmark that assesses SLMs' knowledge understanding through pure speech interactions. Our benchmark 1) uniquely maintains speech format for both inputs and outputs, 2) evaluates model robustness across diverse input audio conditions, and 3) pioneers the assessment of complex tasks like mathematical reasoning in spoken format. Systematic evaluation demonstrates that VoxEval presents significant challenges to current SLMs, revealing their sensitivity to varying audio conditions and highlighting the need to enhance reasoning capabilities in future development. We hope this benchmark could guide the advancement of more sophisticated and reliable SLMs. VoxEval dataset is available at: https://github.com/dreamtheater123/VoxEval
title VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.04962