RAS: a Reliability Oriented Metric for Automatic Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918470607699968 |
|---|---|
| author | Huang, Wenbin Qiu, Yuhang Li, Bohan Guo, Yiwei Peng, Jing Wang, Hankun Chen, Xie Yu, Kai |
| author_facet | Huang, Wenbin Qiu, Yuhang Li, Bohan Guo, Yiwei Peng, Jing Wang, Hankun Chen, Xie Yu, Kai |
| contents | Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream applications. Standard evaluation based on Word Error Rate focuses solely on accuracy and fails to capture transcription reliability. We introduce an abstention-aware transcription framework that enables ASR models to explicitly abstain from uncertain segments. To evaluate reliability under abstention, we propose RAS, a reliability-oriented metric that balances transcription informativeness and error aversion, with its trade-off parameter calibrated by human preference. We then train an abstention-aware ASR model through supervised bootstrapping followed by reinforcement learning. Our experiments demonstrate substantial improvements in transcription reliability while maintaining competitive accuracy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_24278 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | RAS: a Reliability Oriented Metric for Automatic Speech Recognition Huang, Wenbin Qiu, Yuhang Li, Bohan Guo, Yiwei Peng, Jing Wang, Hankun Chen, Xie Yu, Kai Sound Artificial Intelligence Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream applications. Standard evaluation based on Word Error Rate focuses solely on accuracy and fails to capture transcription reliability. We introduce an abstention-aware transcription framework that enables ASR models to explicitly abstain from uncertain segments. To evaluate reliability under abstention, we propose RAS, a reliability-oriented metric that balances transcription informativeness and error aversion, with its trade-off parameter calibrated by human preference. We then train an abstention-aware ASR model through supervised bootstrapping followed by reinforcement learning. Our experiments demonstrate substantial improvements in transcription reliability while maintaining competitive accuracy. |
| title | RAS: a Reliability Oriented Metric for Automatic Speech Recognition |
| topic | Sound Artificial Intelligence |
| url | https://arxiv.org/abs/2604.24278 |