The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gauder, Lara, Riera, Pablo, Slachevsky, Andrea, Forno, Gonzalo, Garcia, Adolfo M., Ferrer, Luciana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
A Toolkit for Detecting Spurious Correlations in Speech Datasets
von: Gauder, Lara, et al.
Veröffentlicht: (2026)
von: Gauder, Lara, et al.
Veröffentlicht: (2026)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
Study on the Fairness of Speaker Verification Systems on Underrepresented Accents in English
von: Estevez, Mariel, et al.
Veröffentlicht: (2022)
von: Estevez, Mariel, et al.
Veröffentlicht: (2022)
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
von: Yang, Bing, et al.
Veröffentlicht: (2024)
von: Yang, Bing, et al.
Veröffentlicht: (2024)
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
Robust Audio Tagging under Class-wise Supervision Unreliability
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
Abusive Speech Detection in Indic Languages Using Acoustic Features
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
von: Spiesberger, Anika A., et al.
Veröffentlicht: (2024)
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
von: Garg, Ashi, et al.
Veröffentlicht: (2025)
von: Garg, Ashi, et al.
Veröffentlicht: (2025)
Improving Acoustic Scene Classification in Low-Resource Conditions
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
Room Impulse Response Generation Conditioned on Acoustic Parameters
von: Arellano, Silvia, et al.
Veröffentlicht: (2025)
von: Arellano, Silvia, et al.
Veröffentlicht: (2025)
DIFFRENT: A Diffusion Model for Recording Environment Transfer of Speech
von: Im, Jaekwon, et al.
Veröffentlicht: (2024)
von: Im, Jaekwon, et al.
Veröffentlicht: (2024)
Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech
von: Daeglau, Mareike, et al.
Veröffentlicht: (2025)
von: Daeglau, Mareike, et al.
Veröffentlicht: (2025)
From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
von: Huang, Mengcheng, et al.
Veröffentlicht: (2026)
von: Huang, Mengcheng, et al.
Veröffentlicht: (2026)
Temporally Heterogeneous Graph Contrastive Learning for Multimodal Acoustic event Classification
von: Chen, Yuanjian, et al.
Veröffentlicht: (2025)
von: Chen, Yuanjian, et al.
Veröffentlicht: (2025)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
Comparative Evaluation of Acoustic Feature Extraction Tools for Clinical Speech Analysis
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
UrBAN: Urban Beehive Acoustics and PheNotyping Dataset
von: Abdollahi, Mahsa, et al.
Veröffentlicht: (2024)
von: Abdollahi, Mahsa, et al.
Veröffentlicht: (2024)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
von: Kirdey, Stanislav
Veröffentlicht: (2025)
von: Kirdey, Stanislav
Veröffentlicht: (2025)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
Evaluating Parkinson's Disease Detection in Anonymized Speech: A Performance and Acoustic Analysis
von: Franzreb, Carlos, et al.
Veröffentlicht: (2026)
von: Franzreb, Carlos, et al.
Veröffentlicht: (2026)
Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings
von: Zhao, He, et al.
Veröffentlicht: (2024)
von: Zhao, He, et al.
Veröffentlicht: (2024)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
Benchmarking Time-localized Explanations for Audio Classification Models
von: Bolaños, Cecilia, et al.
Veröffentlicht: (2025)
von: Bolaños, Cecilia, et al.
Veröffentlicht: (2025)
SEABAD: A Tropical Bird Activity Detection Dataset for Passive Acoustic Monitoring
von: Zabidi, Muhammad Mun'im Ahmad, et al.
Veröffentlicht: (2026)
von: Zabidi, Muhammad Mun'im Ahmad, et al.
Veröffentlicht: (2026)
Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification
von: Gao, Yifan, et al.
Veröffentlicht: (2024)
von: Gao, Yifan, et al.
Veröffentlicht: (2024)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
Improving Design of Input Condition Invariant Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
CUEMPATHY: A Counseling Speech Dataset for Psychotherapy Research
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2026)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2026)
Enforcing Speech Content Privacy in Environmental Sound Recordings using Segment-wise Waveform Reversal
von: Tailleur, Modan, et al.
Veröffentlicht: (2025)
von: Tailleur, Modan, et al.
Veröffentlicht: (2025)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
Effects of Recording Condition and Number of Monitored Days on Discriminative Power of the Daily Phonotrauma Index
von: Ghasemzadeh, Hamzeh, et al.
Veröffentlicht: (2024)
von: Ghasemzadeh, Hamzeh, et al.
Veröffentlicht: (2024)
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
von: Richter, Julius, et al.
Veröffentlicht: (2025)
von: Richter, Julius, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023) -
A Toolkit for Detecting Spurious Correlations in Speech Datasets
von: Gauder, Lara, et al.
Veröffentlicht: (2026) -
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024) -
Study on the Fairness of Speaker Verification Systems on Underrepresented Accents in English
von: Estevez, Mariel, et al.
Veröffentlicht: (2022) -
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
von: Yang, Bing, et al.
Veröffentlicht: (2024)