Back to Basics: Revisiting ASR in the Age of Voice Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tay, Geeyang, Ma, Wentao, Lee, Jaewon, Tang, Yuzhi, Lee, Daniel, Yin, Weisu, Shen, Dongming, Meng, Silin, Zhu, Yi, Li, Mu, Smola, Alex
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912983987257344
author Tay, Geeyang
Ma, Wentao
Lee, Jaewon
Tang, Yuzhi
Lee, Daniel
Yin, Weisu
Shen, Dongming
Meng, Silin
Zhu, Yi
Li, Mu
Smola, Alex
author_facet Tay, Geeyang
Ma, Wentao
Lee, Jaewon
Tang, Yuzhi
Lee, Daniel
Yin, Weisu
Shen, Dongming
Meng, Silin
Zhu, Yi
Li, Mu
Smola, Alex
contents Automatic speech recognition (ASR) systems have achieved near-human accuracy on curated benchmarks, yet still fail in real-world voice agents under conditions that current evaluations do not systematically cover. Without diagnostic tools that isolate specific failure factors, practitioners cannot anticipate which conditions, in which languages, will cause what degree of degradation. We introduce WildASR, a multilingual (four-language) diagnostic benchmark sourced entirely from real human speech that factorizes ASR robustness along three axes: environmental degradation, demographic shift, and linguistic diversity. Evaluating seven widely used ASR systems, we find severe and uneven performance degradation, and model robustness does not transfer across languages or conditions. Critically, models often hallucinate plausible but unspoken content under partial or degraded inputs, creating concrete safety risks for downstream agent behavior. Our results demonstrate that targeted, factor-isolated evaluation is essential for understanding and improving ASR reliability in production systems. Besides the benchmark itself, we also present three analytical tools that practitioners can use to guide deployment decisions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25727
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Back to Basics: Revisiting ASR in the Age of Voice Agents
Tay, Geeyang
Ma, Wentao
Lee, Jaewon
Tang, Yuzhi
Lee, Daniel
Yin, Weisu
Shen, Dongming
Meng, Silin
Zhu, Yi
Li, Mu
Smola, Alex
Artificial Intelligence
Multimedia
Automatic speech recognition (ASR) systems have achieved near-human accuracy on curated benchmarks, yet still fail in real-world voice agents under conditions that current evaluations do not systematically cover. Without diagnostic tools that isolate specific failure factors, practitioners cannot anticipate which conditions, in which languages, will cause what degree of degradation. We introduce WildASR, a multilingual (four-language) diagnostic benchmark sourced entirely from real human speech that factorizes ASR robustness along three axes: environmental degradation, demographic shift, and linguistic diversity. Evaluating seven widely used ASR systems, we find severe and uneven performance degradation, and model robustness does not transfer across languages or conditions. Critically, models often hallucinate plausible but unspoken content under partial or degraded inputs, creating concrete safety risks for downstream agent behavior. Our results demonstrate that targeted, factor-isolated evaluation is essential for understanding and improving ASR reliability in production systems. Besides the benchmark itself, we also present three analytical tools that practitioners can use to guide deployment decisions.
title Back to Basics: Revisiting ASR in the Age of Voice Agents
topic Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2603.25727