A Toolkit for Detecting Spurious Correlations in Speech Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Gauder, Lara, Riera, Pablo, Slachevsky, Andrea, Forno, Gonzalo, García, Adolfo M., Ferrer, Luciana |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
by: Gauder, Lara, et al.
Published: (2024)
by: Gauder, Lara, et al.
Published: (2024)
DDFAD: Dataset Distillation Framework for Audio Data
by: Jiang, Wenbo, et al.
Published: (2024)
by: Jiang, Wenbo, et al.
Published: (2024)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
by: Riera, Pablo, et al.
Published: (2026)
by: Riera, Pablo, et al.
Published: (2026)
Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language
by: Wiafe, Isaac, et al.
Published: (2026)
by: Wiafe, Isaac, et al.
Published: (2026)
A Comparison of Speech Data Augmentation Methods Using S3PRL Toolkit
by: Huh, Mina, et al.
Published: (2023)
by: Huh, Mina, et al.
Published: (2023)
Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
by: Xie, Yuankun, et al.
Published: (2025)
by: Xie, Yuankun, et al.
Published: (2025)
Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues
by: Ba, Zhongjie, et al.
Published: (2026)
by: Ba, Zhongjie, et al.
Published: (2026)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
by: Huang, Kuan-Tang, et al.
Published: (2026)
by: Huang, Kuan-Tang, et al.
Published: (2026)
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
by: Huang, Jia-Hong, et al.
Published: (2026)
by: Huang, Jia-Hong, et al.
Published: (2026)
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Dynamic Fusion Multimodal Network for SpeechWellness Detection
by: Sun, Wenqiang, et al.
Published: (2025)
by: Sun, Wenqiang, et al.
Published: (2025)
SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection
by: Jung, Kyudan, et al.
Published: (2026)
by: Jung, Kyudan, et al.
Published: (2026)
The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
by: Bibbó, Gabriel, et al.
Published: (2024)
by: Bibbó, Gabriel, et al.
Published: (2024)
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
by: Dvirniak, Artem, et al.
Published: (2026)
by: Dvirniak, Artem, et al.
Published: (2026)
Emotion Detection in Speech Using Lightweight and Transformer-Based Models: A Comparative and Ablation Study
by: Onyekwelu-Udoka, Lucky, et al.
Published: (2025)
by: Onyekwelu-Udoka, Lucky, et al.
Published: (2025)
Towards Explicit Acoustic Evidence Perception in Audio LLMs for Speech Deepfake Detection
by: Guo, Xiaoxuan, et al.
Published: (2026)
by: Guo, Xiaoxuan, et al.
Published: (2026)
ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement
by: Chhaglani, Bhawana, et al.
Published: (2025)
by: Chhaglani, Bhawana, et al.
Published: (2025)
Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
by: Truong, Duc-Tuan, et al.
Published: (2025)
by: Truong, Duc-Tuan, et al.
Published: (2025)
Generalizable Speech Deepfake Detection via Information Bottleneck Enhanced Adversarial Alignment
by: Huang, Pu, et al.
Published: (2025)
by: Huang, Pu, et al.
Published: (2025)
A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators
by: Pham, Lam, et al.
Published: (2026)
by: Pham, Lam, et al.
Published: (2026)
Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
DDSP-QbE++: Improving Speech Quality for Speech Anonymisation for Atypical Speech
by: Ghosh, Suhita, et al.
Published: (2026)
by: Ghosh, Suhita, et al.
Published: (2026)
Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling
by: Li, Xingyuan, et al.
Published: (2026)
by: Li, Xingyuan, et al.
Published: (2026)
QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection
by: Truong, Duc-Tuan, et al.
Published: (2025)
by: Truong, Duc-Tuan, et al.
Published: (2025)
BDI-Kit Demo: A Toolkit for Programmable and Conversational Data Harmonization
by: Lopez, Roque, et al.
Published: (2026)
by: Lopez, Roque, et al.
Published: (2026)
Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
by: Ji, Zhoulin, et al.
Published: (2024)
by: Ji, Zhoulin, et al.
Published: (2024)
GSRM: Generative Speech Reward Model for Speech RLHF
by: Shen, Maohao, et al.
Published: (2026)
by: Shen, Maohao, et al.
Published: (2026)
SCDF: A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis
by: Staněk, Vojtěch, et al.
Published: (2025)
by: Staněk, Vojtěch, et al.
Published: (2025)
Understanding Frechet Speech Distance for Synthetic Speech Quality Evaluation
by: Kim, June-Woo, et al.
Published: (2026)
by: Kim, June-Woo, et al.
Published: (2026)
SLM-SS: Speech Language Model for Generative Speech Separation
by: Li, Tianhua, et al.
Published: (2026)
by: Li, Tianhua, et al.
Published: (2026)
SpeechQualityLLM: LLM-Based Multimodal Assessment of Speech Quality
by: Monjur, Mahathir, et al.
Published: (2025)
by: Monjur, Mahathir, et al.
Published: (2025)
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
by: Wang, Kaidi, et al.
Published: (2025)
by: Wang, Kaidi, et al.
Published: (2025)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
by: Cheng, Sitong, et al.
Published: (2025)
by: Cheng, Sitong, et al.
Published: (2025)
Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training
by: Zhou, Naisong, et al.
Published: (2025)
by: Zhou, Naisong, et al.
Published: (2025)
Speech Separation for Hearing-Impaired Children in the Classroom
by: Olalere, Feyisayo, et al.
Published: (2025)
by: Olalere, Feyisayo, et al.
Published: (2025)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
by: Sheng, Zhengyan, et al.
Published: (2025)
by: Sheng, Zhengyan, et al.
Published: (2025)
MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement
by: Yu, Xinyue, et al.
Published: (2025)
by: Yu, Xinyue, et al.
Published: (2025)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
by: Ma, Ziyang, et al.
Published: (2024)
by: Ma, Ziyang, et al.
Published: (2024)
Overview of the Amphion Toolkit (v0.2)
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
Similar Items
-
The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
by: Gauder, Lara, et al.
Published: (2024) -
DDFAD: Dataset Distillation Framework for Audio Data
by: Jiang, Wenbo, et al.
Published: (2024) -
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
by: Riera, Pablo, et al.
Published: (2026) -
Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language
by: Wiafe, Isaac, et al.
Published: (2026) -
A Comparison of Speech Data Augmentation Methods Using S3PRL Toolkit
by: Huh, Mina, et al.
Published: (2023)