Salvato in:
Dettagli Bibliografici
Autori principali: Ellis, Zachary, Joselowitz, Jared, Deo, Yash, He, Yajie, Kalygina, Anna, Higham, Aisling, Rahimzadeh, Mana, Jia, Yan, Habli, Ibrahim, Lim, Ernest
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2511.16544
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914262486614016
author Ellis, Zachary
Joselowitz, Jared
Deo, Yash
He, Yajie
Kalygina, Anna
Higham, Aisling
Rahimzadeh, Mana
Jia, Yan
Habli, Ibrahim
Lim, Ernest
author_facet Ellis, Zachary
Joselowitz, Jared
Deo, Yash
He, Yajie
Kalygina, Anna
Higham, Aisling
Rahimzadeh, Mana
Jia, Yan
Habli, Ibrahim
Lim, Ernest
contents As Automatic Speech Recognition (ASR) is increasingly deployed in clinical dialogue, standard evaluations still rely heavily on Word Error Rate (WER). This paper challenges that standard, investigating whether WER or other common metrics correlate with the clinical impact of transcription errors. We establish a gold-standard benchmark by having expert clinicians compare ground-truth utterances to their ASR-generated counterparts, labeling the clinical impact of any discrepancies found in two distinct doctor-patient dialogue datasets. Our analysis reveals that WER and a comprehensive suite of existing metrics correlate poorly with the clinician-assigned risk labels (No, Minimal, or Significant Impact). To bridge this evaluation gap, we introduce an LLM-as-a-Judge, programmatically optimized using GEPA through DSPy to replicate expert clinical assessment. The optimized judge (Gemini-2.5-Pro) achieves human-comparable performance, obtaining 90% accuracy and a strong Cohen's kappa of 0.816. This work provides a validated, automated framework for moving ASR evaluation beyond simple textual fidelity to a necessary, scalable assessment of safety in clinical dialogue.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16544
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue
Ellis, Zachary
Joselowitz, Jared
Deo, Yash
He, Yajie
Kalygina, Anna
Higham, Aisling
Rahimzadeh, Mana
Jia, Yan
Habli, Ibrahim
Lim, Ernest
Computation and Language
Artificial Intelligence
As Automatic Speech Recognition (ASR) is increasingly deployed in clinical dialogue, standard evaluations still rely heavily on Word Error Rate (WER). This paper challenges that standard, investigating whether WER or other common metrics correlate with the clinical impact of transcription errors. We establish a gold-standard benchmark by having expert clinicians compare ground-truth utterances to their ASR-generated counterparts, labeling the clinical impact of any discrepancies found in two distinct doctor-patient dialogue datasets. Our analysis reveals that WER and a comprehensive suite of existing metrics correlate poorly with the clinician-assigned risk labels (No, Minimal, or Significant Impact). To bridge this evaluation gap, we introduce an LLM-as-a-Judge, programmatically optimized using GEPA through DSPy to replicate expert clinical assessment. The optimized judge (Gemini-2.5-Pro) achieves human-comparable performance, obtaining 90% accuracy and a strong Cohen's kappa of 0.816. This work provides a validated, automated framework for moving ASR evaluation beyond simple textual fidelity to a necessary, scalable assessment of safety in clinical dialogue.
title WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.16544