A Tutorial on Clinical Speech AI Development: From Data Collection to Model Validation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ng, Si-Ioi, Xu, Lingfeng, Siegert, Ingo, Cummins, Nicholas, Benway, Nina R., Liss, Julie, Berisha, Visar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
Automated Extraction of Spatio-Semantic Graphs for Identifying Cognitive Impairment
von: Ng, Si-Ioi, et al.
Veröffentlicht: (2025)
von: Ng, Si-Ioi, et al.
Veröffentlicht: (2025)
Requirements for Mass Adoption of Assistive Listening Technology by the General Public
von: Kaufmann, Thomas B., et al.
Veröffentlicht: (2023)
von: Kaufmann, Thomas B., et al.
Veröffentlicht: (2023)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
Advancing Automated Spatio-Semantic Analysis in Picture Description Using Language Models
von: Ng, Si-Ioi, et al.
Veröffentlicht: (2025)
von: Ng, Si-Ioi, et al.
Veröffentlicht: (2025)
Why Speech Deepfake Detectors Won't Generalize: The Limits of Detection in an Open World
von: Berisha, Visar, et al.
Veröffentlicht: (2025)
von: Berisha, Visar, et al.
Veröffentlicht: (2025)
A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
Towards robust paralinguistic assessment for real-world mobile health (mHealth) monitoring: an initial study of reverberation effects on speech
von: Dineley, Judith, et al.
Veröffentlicht: (2023)
von: Dineley, Judith, et al.
Veröffentlicht: (2023)
TidyVoice 2026 Challenge Evaluation Plan
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
Reference-free Adversarial Sex Obfuscation in Speech
von: Qu, Yangyang, et al.
Veröffentlicht: (2025)
von: Qu, Yangyang, et al.
Veröffentlicht: (2025)
Wireless Hearables With Programmable Speech AI Accelerators
von: Itani, Malek, et al.
Veröffentlicht: (2025)
von: Itani, Malek, et al.
Veröffentlicht: (2025)
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis
von: R, Vinotha, et al.
Veröffentlicht: (2024)
von: R, Vinotha, et al.
Veröffentlicht: (2024)
Comparative Evaluation of Acoustic Feature Extraction Tools for Clinical Speech Analysis
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
Steered Response Power for Sound Source Localization: A Tutorial Review
von: Grinstein, Eric, et al.
Veröffentlicht: (2024)
von: Grinstein, Eric, et al.
Veröffentlicht: (2024)
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
von: Garg, Ashi, et al.
Veröffentlicht: (2025)
von: Garg, Ashi, et al.
Veröffentlicht: (2025)
From Sharpness to Better Generalization for Speech Deepfake Detection
von: Huang, Wen, et al.
Veröffentlicht: (2025)
von: Huang, Wen, et al.
Veröffentlicht: (2025)
From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
von: Huang, Mengcheng, et al.
Veröffentlicht: (2026)
von: Huang, Mengcheng, et al.
Veröffentlicht: (2026)
AntiDeepFake: AI for Deep Fake Speech Recognition
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2024)
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2024)
Universal Speech Content Factorization
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2026)
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2026)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
Application of Whisper in Clinical Practice: the Post-Stroke Speech Assessment during a Naming Task
von: Davudova, Milena, et al.
Veröffentlicht: (2025)
von: Davudova, Milena, et al.
Veröffentlicht: (2025)
On the Relevance of Clinical Assessment Tasks for the Automatic Detection of Parkinson's Disease Medication State from Speech
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2025)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2025)
Anonymising Elderly and Pathological Speech: Voice Conversion Using DDSP and Query-by-Example
von: Ghosh, Suhita, et al.
Veröffentlicht: (2024)
von: Ghosh, Suhita, et al.
Veröffentlicht: (2024)
Best Practices and Considerations for Child Speech Corpus Collection and Curation in Educational, Clinical, and Forensic Scenarios
von: Hansen, John, et al.
Veröffentlicht: (2025)
von: Hansen, John, et al.
Veröffentlicht: (2025)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline
von: Raghav, Nikhil
Veröffentlicht: (2026)
von: Raghav, Nikhil
Veröffentlicht: (2026)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2025)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2025)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
von: Nasretdinov, Rauf, et al.
Veröffentlicht: (2025)
von: Nasretdinov, Rauf, et al.
Veröffentlicht: (2025)
A Neural Speech Codec for Noise Robust Speech Coding
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025) -
Automated Extraction of Spatio-Semantic Graphs for Identifying Cognitive Impairment
von: Ng, Si-Ioi, et al.
Veröffentlicht: (2025) -
Requirements for Mass Adoption of Assistive Listening Technology by the General Public
von: Kaufmann, Thomas B., et al.
Veröffentlicht: (2023) -
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026) -
Advancing Automated Spatio-Semantic Analysis in Picture Description Using Language Models
von: Ng, Si-Ioi, et al.
Veröffentlicht: (2025)