Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Rosero, Karen, Yeo, Eunjung, Mortensen, David R., Slot, Cortney Van't, Hallac, Rami R., Busso, Carlos |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
por: Yeo, Eunjung
Publicado: (2024)
por: Yeo, Eunjung
Publicado: (2024)
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
por: Yeo, Eunjung, et al.
Publicado: (2025)
por: Yeo, Eunjung, et al.
Publicado: (2025)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
por: Yeo, Eunjung, et al.
Publicado: (2026)
por: Yeo, Eunjung, et al.
Publicado: (2026)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
por: Choi, Kwanghee, et al.
Publicado: (2026)
por: Choi, Kwanghee, et al.
Publicado: (2026)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
por: Choi, Kwanghee, et al.
Publicado: (2026)
por: Choi, Kwanghee, et al.
Publicado: (2026)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
por: Li, Chin-Jou, et al.
Publicado: (2025)
por: Li, Chin-Jou, et al.
Publicado: (2025)
PRiSM: Benchmarking Phone Realization in Speech Models
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
An Empirical Recipe for Universal Phone Recognition
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
Improving Speech Emotion Recognition with Mutual Information Regularized Generative Model
por: Ahn, Chung-Soo, et al.
Publicado: (2025)
por: Ahn, Chung-Soo, et al.
Publicado: (2025)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
por: Lau, Hok-Shing, et al.
Publicado: (2024)
por: Lau, Hok-Shing, et al.
Publicado: (2024)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
por: Kim, Taesoo, et al.
Publicado: (2025)
por: Kim, Taesoo, et al.
Publicado: (2025)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
por: Sun, Haitong, et al.
Publicado: (2026)
por: Sun, Haitong, et al.
Publicado: (2026)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
por: Huo, Mingyue, et al.
Publicado: (2025)
por: Huo, Mingyue, et al.
Publicado: (2025)
CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
por: Su, Xiaosu, et al.
Publicado: (2026)
por: Su, Xiaosu, et al.
Publicado: (2026)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
por: Suda, Hitoshi, et al.
Publicado: (2025)
por: Suda, Hitoshi, et al.
Publicado: (2025)
A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
por: Upadhyay, Shreya G., et al.
Publicado: (2024)
por: Upadhyay, Shreya G., et al.
Publicado: (2024)
Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
por: Sun, Esther, et al.
Publicado: (2026)
por: Sun, Esther, et al.
Publicado: (2026)
RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
por: Bargum, Anders R., et al.
Publicado: (2024)
por: Bargum, Anders R., et al.
Publicado: (2024)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
por: Du, Zongyang, et al.
Publicado: (2025)
por: Du, Zongyang, et al.
Publicado: (2025)
MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors
por: Bao, Guangyin, et al.
Publicado: (2026)
por: Bao, Guangyin, et al.
Publicado: (2026)
Adapting Speech Language Model to Singing Voice Synthesis
por: Zhao, Yiwen, et al.
Publicado: (2025)
por: Zhao, Yiwen, et al.
Publicado: (2025)
An Agent-Based Framework for Automated Higher-Voice Harmony Generation
por: Ganapathy, Nia D'Souza, et al.
Publicado: (2025)
por: Ganapathy, Nia D'Souza, et al.
Publicado: (2025)
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
por: Zhou, Yixuan, et al.
Publicado: (2025)
por: Zhou, Yixuan, et al.
Publicado: (2025)
SpeechT: Findings of the First Mentorship in Speech Translation
por: Moslem, Yasmin, et al.
Publicado: (2025)
por: Moslem, Yasmin, et al.
Publicado: (2025)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
por: Ulgen, Ismail Rasim, et al.
Publicado: (2024)
por: Ulgen, Ismail Rasim, et al.
Publicado: (2024)
Cross-Technology Generalization in Synthesized Speech Detection: Evaluating AST Models with Modern Voice Generators
por: Ustinov, Andrew, et al.
Publicado: (2025)
por: Ustinov, Andrew, et al.
Publicado: (2025)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
por: Hernandez, Abner, et al.
Publicado: (2026)
por: Hernandez, Abner, et al.
Publicado: (2026)
Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
por: Tian, Yusheng, et al.
Publicado: (2024)
por: Tian, Yusheng, et al.
Publicado: (2024)
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
por: Cheng, Changhao, et al.
Publicado: (2026)
por: Cheng, Changhao, et al.
Publicado: (2026)
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
por: Xue, Jun, et al.
Publicado: (2026)
por: Xue, Jun, et al.
Publicado: (2026)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
por: Cheng, Sitong, et al.
Publicado: (2025)
por: Cheng, Sitong, et al.
Publicado: (2025)
A Probabilistic Generative Model for Spectral Speech Enhancement
por: Hidalgo-Araya, Marco, et al.
Publicado: (2026)
por: Hidalgo-Araya, Marco, et al.
Publicado: (2026)
Speech to Speech Synthesis for Voice Impersonation
por: Johnson, Bjorn, et al.
Publicado: (2026)
por: Johnson, Bjorn, et al.
Publicado: (2026)
Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
por: Upadhyay, Shreya G., et al.
Publicado: (2024)
por: Upadhyay, Shreya G., et al.
Publicado: (2024)
Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
por: Kunze, Tarek, et al.
Publicado: (2025)
por: Kunze, Tarek, et al.
Publicado: (2025)
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
por: Lee, Jaejun, et al.
Publicado: (2026)
por: Lee, Jaejun, et al.
Publicado: (2026)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
por: Zheng, Zhisheng, et al.
Publicado: (2025)
por: Zheng, Zhisheng, et al.
Publicado: (2025)
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
por: Li, Xueyan, et al.
Publicado: (2025)
por: Li, Xueyan, et al.
Publicado: (2025)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
por: Byun, Kyungguen, et al.
Publicado: (2025)
por: Byun, Kyungguen, et al.
Publicado: (2025)
Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
por: Liu, Qingyu, et al.
Publicado: (2025)
por: Liu, Qingyu, et al.
Publicado: (2025)
Ejemplares similares
-
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
por: Yeo, Eunjung
Publicado: (2024) -
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
por: Yeo, Eunjung, et al.
Publicado: (2025) -
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
por: Yeo, Eunjung, et al.
Publicado: (2026) -
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
por: Choi, Kwanghee, et al.
Publicado: (2026) -
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
por: Choi, Kwanghee, et al.
Publicado: (2026)