Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909543076724736 |
|---|---|
| author | Kuhn, Korbinian Kersken, Verena Zimmermann, Gottfried |
| author_facet | Kuhn, Korbinian Kersken, Verena Zimmermann, Gottfried |
| contents | Despite advances in Automatic Speech Recognition (ASR), transcription errors persist and require manual correction. Confidence scores, which indicate the certainty of ASR results, could assist users in identifying and correcting errors. This study evaluates the reliability of confidence scores for error detection through a comprehensive analysis of end-to-end ASR models and a user study with 36 participants. The results show that while confidence scores correlate with transcription accuracy, their error detection performance is limited. Classifiers frequently miss errors or generate many false positives, undermining their practical utility. Confidence-based error detection neither improved correction efficiency nor was perceived as helpful by participants. These findings highlight the limitations of confidence scores and the need for more sophisticated approaches to improve user interaction and explainability of ASR results. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_15124 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces Kuhn, Korbinian Kersken, Verena Zimmermann, Gottfried Human-Computer Interaction Computation and Language Sound Audio and Speech Processing I.2.7 Despite advances in Automatic Speech Recognition (ASR), transcription errors persist and require manual correction. Confidence scores, which indicate the certainty of ASR results, could assist users in identifying and correcting errors. This study evaluates the reliability of confidence scores for error detection through a comprehensive analysis of end-to-end ASR models and a user study with 36 participants. The results show that while confidence scores correlate with transcription accuracy, their error detection performance is limited. Classifiers frequently miss errors or generate many false positives, undermining their practical utility. Confidence-based error detection neither improved correction efficiency nor was perceived as helpful by participants. These findings highlight the limitations of confidence scores and the need for more sophisticated approaches to improve user interaction and explainability of ASR results. |
| title | Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces |
| topic | Human-Computer Interaction Computation and Language Sound Audio and Speech Processing I.2.7 |
| url | https://arxiv.org/abs/2503.15124 |