Evaluating Speech-to-Text Systems with PennSound

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wright, Jonathan, Liberman, Mark, Ryant, Neville, Fiumara, James
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908306618974208
author Wright, Jonathan
Liberman, Mark
Ryant, Neville
Fiumara, James
author_facet Wright, Jonathan
Liberman, Mark
Ryant, Neville
Fiumara, James
contents A random sample of nearly 10 hours of speech from PennSound, the world's largest online collection of poetry readings and discussions, was used as a benchmark to evaluate several commercial and open-source speech-to-text systems. PennSound's wide variation in recording conditions and speech styles makes it a good representative for many other untranscribed audio collections. Reference transcripts were created by trained annotators, and system transcripts were produced from AWS, Azure, Google, IBM, NeMo, Rev.ai, Whisper, and Whisper.cpp. Based on word error rate, Rev.ai was the top performer, and Whisper was the top open source performer (as long as hallucinations were avoided). AWS had the best diarization error rates among three systems. However, WER and DER differences were slim, and various tradeoffs may motivate choosing different systems for different end users. We also examine the issue of hallucinations in Whisper. Users of Whisper should be cautioned to be aware of runtime options, and whether the speed vs accuracy trade off is acceptable.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05702
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Speech-to-Text Systems with PennSound
Wright, Jonathan
Liberman, Mark
Ryant, Neville
Fiumara, James
Computation and Language
A random sample of nearly 10 hours of speech from PennSound, the world's largest online collection of poetry readings and discussions, was used as a benchmark to evaluate several commercial and open-source speech-to-text systems. PennSound's wide variation in recording conditions and speech styles makes it a good representative for many other untranscribed audio collections. Reference transcripts were created by trained annotators, and system transcripts were produced from AWS, Azure, Google, IBM, NeMo, Rev.ai, Whisper, and Whisper.cpp. Based on word error rate, Rev.ai was the top performer, and Whisper was the top open source performer (as long as hallucinations were avoided). AWS had the best diarization error rates among three systems. However, WER and DER differences were slim, and various tradeoffs may motivate choosing different systems for different end users. We also examine the issue of hallucinations in Whisper. Users of Whisper should be cautioned to be aware of runtime options, and whether the speed vs accuracy trade off is acceptable.
title Evaluating Speech-to-Text Systems with PennSound
topic Computation and Language
url https://arxiv.org/abs/2504.05702