Beyond Word Error Rate: Auditing the Diversity Tax in Speech Recognition through Dataset Cartography

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Ting-Hui, Clemmensen, Line H., Das, Sneha
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911489157234688
author Cheng, Ting-Hui
Clemmensen, Line H.
Das, Sneha
author_facet Cheng, Ting-Hui
Clemmensen, Line H.
Das, Sneha
contents Automatic speech recognition (ASR) systems are predominantly evaluated using the Word Error Rate (WER). However, raw token-level metrics fail to capture semantic fidelity and routinely obscures the `diversity tax', the disproportionate burden on marginalized and atypical speaker due to systematic recognition failures. In this paper, we explore the limitations of relying solely on lexical counts by systematically evaluating a broader class of non-linear and semantic metrics. To enable rigorous model auditing, we introduce the sample difficulty index (SDI), a novel metric that quantifies how intrinsic demographic and acoustic factors drive model failure. By mapping SDI on data cartography, we demonstrate that metrics EmbER and SemDist expose hidden systemic biases and inter-model disagreements that WER ignores. Finally, our findings are the first steps towards a robust audit framework for prospective safety analysis, empowering developers to audit and mitigate ASR disparities prior to deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05267
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Word Error Rate: Auditing the Diversity Tax in Speech Recognition through Dataset Cartography
Cheng, Ting-Hui
Clemmensen, Line H.
Das, Sneha
Machine Learning
Automatic speech recognition (ASR) systems are predominantly evaluated using the Word Error Rate (WER). However, raw token-level metrics fail to capture semantic fidelity and routinely obscures the `diversity tax', the disproportionate burden on marginalized and atypical speaker due to systematic recognition failures. In this paper, we explore the limitations of relying solely on lexical counts by systematically evaluating a broader class of non-linear and semantic metrics. To enable rigorous model auditing, we introduce the sample difficulty index (SDI), a novel metric that quantifies how intrinsic demographic and acoustic factors drive model failure. By mapping SDI on data cartography, we demonstrate that metrics EmbER and SemDist expose hidden systemic biases and inter-model disagreements that WER ignores. Finally, our findings are the first steps towards a robust audit framework for prospective safety analysis, empowering developers to audit and mitigate ASR disparities prior to deployment.
title Beyond Word Error Rate: Auditing the Diversity Tax in Speech Recognition through Dataset Cartography
topic Machine Learning
url https://arxiv.org/abs/2603.05267