A Toolkit for Detecting Spurious Correlations in Speech Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gauder, Lara, Riera, Pablo, Slachevsky, Andrea, Forno, Gonzalo, García, Adolfo M., Ferrer, Luciana
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913072915939328
author Gauder, Lara
Riera, Pablo
Slachevsky, Andrea
Forno, Gonzalo
García, Adolfo M.
Ferrer, Luciana
author_facet Gauder, Lara
Riera, Pablo
Slachevsky, Andrea
Forno, Gonzalo
García, Adolfo M.
Ferrer, Luciana
contents We introduce a toolkit for uncovering spurious correlations between recording characteristics and target class in speech datasets. Spurious correlations may arise due to heterogeneous recording conditions, a common scenario for health-related datasets. When present both in the training and test data, these correlations result in an overestimation of the system performance -- a dangerous situation, specially in high-stakes application where systems are required to satisfy minimum performance requirements. Our toolkit implements a diagnostic method based on the detection of the target class using only the non-speech regions in the audio. Better than chance performance at this task indicates that information about the target class can be extracted from the non-speech regions, flagging the presence of spurious correlations. The toolkit is publicly available for research use.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26676
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Toolkit for Detecting Spurious Correlations in Speech Datasets
Gauder, Lara
Riera, Pablo
Slachevsky, Andrea
Forno, Gonzalo
García, Adolfo M.
Ferrer, Luciana
Sound
Artificial Intelligence
Databases
We introduce a toolkit for uncovering spurious correlations between recording characteristics and target class in speech datasets. Spurious correlations may arise due to heterogeneous recording conditions, a common scenario for health-related datasets. When present both in the training and test data, these correlations result in an overestimation of the system performance -- a dangerous situation, specially in high-stakes application where systems are required to satisfy minimum performance requirements. Our toolkit implements a diagnostic method based on the detection of the target class using only the non-speech regions in the audio. Better than chance performance at this task indicates that information about the target class can be extracted from the non-speech regions, flagging the presence of spurious correlations. The toolkit is publicly available for research use.
title A Toolkit for Detecting Spurious Correlations in Speech Datasets
topic Sound
Artificial Intelligence
Databases
url https://arxiv.org/abs/2604.26676