VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Hongcheng, Hou, Yixuan, Liu, Heyang, Wang, Yuhao, Wang, Yanfeng, Wang, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908599392927744
author Liu, Hongcheng
Hou, Yixuan
Liu, Heyang
Wang, Yuhao
Wang, Yanfeng
Wang, Yu
author_facet Liu, Hongcheng
Hou, Yixuan
Liu, Heyang
Wang, Yuhao
Wang, Yanfeng
Wang, Yu
contents While Speech Large Language Models (Speech-LLMs) show strong performance in many applications, their robustness is critically under-tested, especially to speech disfluency. Existing evaluations often rely on idealized inputs, overlooking common disfluencies, particularly those associated with conditions like Parkinson's disease. This work investigates whether current Speech-LLMs can maintain performance when interacting with users who have speech impairments. To facilitate this inquiry, we introduce VocalBench-DF, a framework for the systematic evaluation of disfluency across a multi-dimensional taxonomy. Our evaluation of 22 mainstream Speech-LLMs reveals substantial performance degradation, indicating that their real-world readiness is limited. Further analysis identifies phoneme-level processing and long-context modeling as primary bottlenecks responsible for these failures. Strengthening recognition and reasoning capability from components and pipelines can substantially improve robustness. These findings highlight the urgent need for new methods to improve disfluency handling and build truly inclusive Speech-LLMs
format Preprint
id arxiv_https___arxiv_org_abs_2510_15406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
Liu, Hongcheng
Hou, Yixuan
Liu, Heyang
Wang, Yuhao
Wang, Yanfeng
Wang, Yu
Computation and Language
While Speech Large Language Models (Speech-LLMs) show strong performance in many applications, their robustness is critically under-tested, especially to speech disfluency. Existing evaluations often rely on idealized inputs, overlooking common disfluencies, particularly those associated with conditions like Parkinson's disease. This work investigates whether current Speech-LLMs can maintain performance when interacting with users who have speech impairments. To facilitate this inquiry, we introduce VocalBench-DF, a framework for the systematic evaluation of disfluency across a multi-dimensional taxonomy. Our evaluation of 22 mainstream Speech-LLMs reveals substantial performance degradation, indicating that their real-world readiness is limited. Further analysis identifies phoneme-level processing and long-context modeling as primary bottlenecks responsible for these failures. Strengthening recognition and reasoning capability from components and pipelines can substantially improve robustness. These findings highlight the urgent need for new methods to improve disfluency handling and build truly inclusive Speech-LLMs
title VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
topic Computation and Language
url https://arxiv.org/abs/2510.15406