RAS: a Reliability Oriented Metric for Automatic Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wenbin, Qiu, Yuhang, Li, Bohan, Guo, Yiwei, Peng, Jing, Wang, Hankun, Chen, Xie, Yu, Kai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918470607699968
author Huang, Wenbin
Qiu, Yuhang
Li, Bohan
Guo, Yiwei
Peng, Jing
Wang, Hankun
Chen, Xie
Yu, Kai
author_facet Huang, Wenbin
Qiu, Yuhang
Li, Bohan
Guo, Yiwei
Peng, Jing
Wang, Hankun
Chen, Xie
Yu, Kai
contents Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream applications. Standard evaluation based on Word Error Rate focuses solely on accuracy and fails to capture transcription reliability. We introduce an abstention-aware transcription framework that enables ASR models to explicitly abstain from uncertain segments. To evaluate reliability under abstention, we propose RAS, a reliability-oriented metric that balances transcription informativeness and error aversion, with its trade-off parameter calibrated by human preference. We then train an abstention-aware ASR model through supervised bootstrapping followed by reinforcement learning. Our experiments demonstrate substantial improvements in transcription reliability while maintaining competitive accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24278
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RAS: a Reliability Oriented Metric for Automatic Speech Recognition
Huang, Wenbin
Qiu, Yuhang
Li, Bohan
Guo, Yiwei
Peng, Jing
Wang, Hankun
Chen, Xie
Yu, Kai
Sound
Artificial Intelligence
Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream applications. Standard evaluation based on Word Error Rate focuses solely on accuracy and fails to capture transcription reliability. We introduce an abstention-aware transcription framework that enables ASR models to explicitly abstain from uncertain segments. To evaluate reliability under abstention, we propose RAS, a reliability-oriented metric that balances transcription informativeness and error aversion, with its trade-off parameter calibrated by human preference. We then train an abstention-aware ASR model through supervised bootstrapping followed by reinforcement learning. Our experiments demonstrate substantial improvements in transcription reliability while maintaining competitive accuracy.
title RAS: a Reliability Oriented Metric for Automatic Speech Recognition
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2604.24278