Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tseng, Yuan, Parcollet, Titouan, van Dalen, Rogier, Zhang, Shucong, Bhattacharya, Sourav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912414142824448
author Tseng, Yuan
Parcollet, Titouan
van Dalen, Rogier
Zhang, Shucong
Bhattacharya, Sourav
author_facet Tseng, Yuan
Parcollet, Titouan
van Dalen, Rogier
Zhang, Shucong
Bhattacharya, Sourav
contents Recent work suggests that large language models (LLMs) can improve performance of speech tasks compared to existing systems. To support their claims, results on LibriSpeech and Common Voice are often quoted. However, this work finds that a substantial amount of the LibriSpeech and Common Voice evaluation sets appear in public LLM pretraining corpora. This calls into question the reliability of findings drawn from these two datasets. To measure contamination impact, LLMs trained with/without contamination are compared. A contaminated LLM is more likely to generate test sentences it has seen during training. Then, speech recognisers based on LLMs are compared. They show only subtle error rate differences if the LLM is contaminated, but assign significantly higher probabilities to transcriptions seen during LLM training. Results show that LLM outputs can be biased by tiny amounts of data contamination, highlighting the importance of evaluating LLM-based speech systems with held-out data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22251
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
Tseng, Yuan
Parcollet, Titouan
van Dalen, Rogier
Zhang, Shucong
Bhattacharya, Sourav
Audio and Speech Processing
Computation and Language
Recent work suggests that large language models (LLMs) can improve performance of speech tasks compared to existing systems. To support their claims, results on LibriSpeech and Common Voice are often quoted. However, this work finds that a substantial amount of the LibriSpeech and Common Voice evaluation sets appear in public LLM pretraining corpora. This calls into question the reliability of findings drawn from these two datasets. To measure contamination impact, LLMs trained with/without contamination are compared. A contaminated LLM is more likely to generate test sentences it has seen during training. Then, speech recognisers based on LLMs are compared. They show only subtle error rate differences if the LLM is contaminated, but assign significantly higher probabilities to transcriptions seen during LLM training. Results show that LLM outputs can be biased by tiny amounts of data contamination, highlighting the importance of evaluating LLM-based speech systems with held-out data.
title Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2505.22251