A Toolkit for Detecting Spurious Correlations in Speech Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913072915939328 |
|---|---|
| author | Gauder, Lara Riera, Pablo Slachevsky, Andrea Forno, Gonzalo García, Adolfo M. Ferrer, Luciana |
| author_facet | Gauder, Lara Riera, Pablo Slachevsky, Andrea Forno, Gonzalo García, Adolfo M. Ferrer, Luciana |
| contents | We introduce a toolkit for uncovering spurious correlations between recording characteristics and target class in speech datasets. Spurious correlations may arise due to heterogeneous recording conditions, a common scenario for health-related datasets. When present both in the training and test data, these correlations result in an overestimation of the system performance -- a dangerous situation, specially in high-stakes application where systems are required to satisfy minimum performance requirements. Our toolkit implements a diagnostic method based on the detection of the target class using only the non-speech regions in the audio. Better than chance performance at this task indicates that information about the target class can be extracted from the non-speech regions, flagging the presence of spurious correlations. The toolkit is publicly available for research use. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_26676 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | A Toolkit for Detecting Spurious Correlations in Speech Datasets Gauder, Lara Riera, Pablo Slachevsky, Andrea Forno, Gonzalo García, Adolfo M. Ferrer, Luciana Sound Artificial Intelligence Databases We introduce a toolkit for uncovering spurious correlations between recording characteristics and target class in speech datasets. Spurious correlations may arise due to heterogeneous recording conditions, a common scenario for health-related datasets. When present both in the training and test data, these correlations result in an overestimation of the system performance -- a dangerous situation, specially in high-stakes application where systems are required to satisfy minimum performance requirements. Our toolkit implements a diagnostic method based on the detection of the target class using only the non-speech regions in the audio. Better than chance performance at this task indicates that information about the target class can be extracted from the non-speech regions, flagging the presence of spurious correlations. The toolkit is publicly available for research use. |
| title | A Toolkit for Detecting Spurious Correlations in Speech Datasets |
| topic | Sound Artificial Intelligence Databases |
| url | https://arxiv.org/abs/2604.26676 |