Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuhn, Korbinian, Kersken, Verena, Zimmermann, Gottfried
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909543076724736
author Kuhn, Korbinian
Kersken, Verena
Zimmermann, Gottfried
author_facet Kuhn, Korbinian
Kersken, Verena
Zimmermann, Gottfried
contents Despite advances in Automatic Speech Recognition (ASR), transcription errors persist and require manual correction. Confidence scores, which indicate the certainty of ASR results, could assist users in identifying and correcting errors. This study evaluates the reliability of confidence scores for error detection through a comprehensive analysis of end-to-end ASR models and a user study with 36 participants. The results show that while confidence scores correlate with transcription accuracy, their error detection performance is limited. Classifiers frequently miss errors or generate many false positives, undermining their practical utility. Confidence-based error detection neither improved correction efficiency nor was perceived as helpful by participants. These findings highlight the limitations of confidence scores and the need for more sophisticated approaches to improve user interaction and explainability of ASR results.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15124
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces
Kuhn, Korbinian
Kersken, Verena
Zimmermann, Gottfried
Human-Computer Interaction
Computation and Language
Sound
Audio and Speech Processing
I.2.7
Despite advances in Automatic Speech Recognition (ASR), transcription errors persist and require manual correction. Confidence scores, which indicate the certainty of ASR results, could assist users in identifying and correcting errors. This study evaluates the reliability of confidence scores for error detection through a comprehensive analysis of end-to-end ASR models and a user study with 36 participants. The results show that while confidence scores correlate with transcription accuracy, their error detection performance is limited. Classifiers frequently miss errors or generate many false positives, undermining their practical utility. Confidence-based error detection neither improved correction efficiency nor was perceived as helpful by participants. These findings highlight the limitations of confidence scores and the need for more sophisticated approaches to improve user interaction and explainability of ASR results.
title Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces
topic Human-Computer Interaction
Computation and Language
Sound
Audio and Speech Processing
I.2.7
url https://arxiv.org/abs/2503.15124