Voice Disorder Analysis: a Transformer-based Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Koudounas, Alkis, Ciravegna, Gabriele, Fantini, Marco, Succo, Giovanni, Crosetti, Erika, Cerquitelli, Tania, Baralis, Elena |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Concept-based approach to Voice Disorder Detection
by: Ghia, Davide, et al.
Published: (2025)
by: Ghia, Davide, et al.
Published: (2025)
MVP: Multi-source Voice Pathology detection
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
A Contrastive Learning Approach to Mitigate Bias in Speech Models
by: Koudounas, Alkis, et al.
Published: (2024)
by: Koudounas, Alkis, et al.
Published: (2024)
"KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models
by: Koudounas, Alkis, et al.
Published: (2024)
by: Koudounas, Alkis, et al.
Published: (2024)
Benchmarking Representations for Speech, Music, and Acoustic Events
by: La Quatra, Moreno, et al.
Published: (2024)
by: La Quatra, Moreno, et al.
Published: (2024)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
voc2vec: A Foundation Model for Non-Verbal Vocalization
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
"Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
Hallucination Benchmark for Speech Foundation Models
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
ITALIC: An Italian Intent Classification Dataset
by: Koudounas, Alkis, et al.
Published: (2023)
by: Koudounas, Alkis, et al.
Published: (2023)
Exploring Generative Error Correction for Dysarthric Speech Recognition
by: La Quatra, Moreno, et al.
Published: (2025)
by: La Quatra, Moreno, et al.
Published: (2025)
Zero-shot Voice Conversion with Diffusion Transformers
by: Liu, Songting
Published: (2024)
by: Liu, Songting
Published: (2024)
VANPY: Voice Analysis Framework
by: Koushnir, Gregory, et al.
Published: (2025)
by: Koushnir, Gregory, et al.
Published: (2025)
A Novel Stochastic Transformer-based Approach for Post-Traumatic Stress Disorder Detection using Audio Recording of Clinical Interviews
by: Dia, Mamadou, et al.
Published: (2024)
by: Dia, Mamadou, et al.
Published: (2024)
OpenVoice: Versatile Instant Voice Cloning
by: Qin, Zengyi, et al.
Published: (2023)
by: Qin, Zengyi, et al.
Published: (2023)
HLTCOE JHU Submission to the Voice Privacy Challenge 2024
by: Xinyuan, Henry Li, et al.
Published: (2024)
by: Xinyuan, Henry Li, et al.
Published: (2024)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
by: Lee, Philip H., et al.
Published: (2024)
by: Lee, Philip H., et al.
Published: (2024)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
by: Morrone, Giovanni, et al.
Published: (2023)
by: Morrone, Giovanni, et al.
Published: (2023)
GenVC: Self-Supervised Zero-Shot Voice Conversion
by: Cai, Zexin, et al.
Published: (2025)
by: Cai, Zexin, et al.
Published: (2025)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
by: Du, Zongyang, et al.
Published: (2025)
by: Du, Zongyang, et al.
Published: (2025)
Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
by: Ariyanti, Whenty, et al.
Published: (2025)
by: Ariyanti, Whenty, et al.
Published: (2025)
Voice-Driven Mortality Prediction in Hospitalized Heart Failure Patients: A Machine Learning Approach Enhanced with Diagnostic Biomarkers
by: Ahmadli, Nihat, et al.
Published: (2024)
by: Ahmadli, Nihat, et al.
Published: (2024)
Compact Neural TTS Voices for Accessibility
by: Jain, Kunal, et al.
Published: (2025)
by: Jain, Kunal, et al.
Published: (2025)
Speech to Speech Synthesis for Voice Impersonation
by: Johnson, Bjorn, et al.
Published: (2026)
by: Johnson, Bjorn, et al.
Published: (2026)
Discrete Optimal Transport and Voice Conversion
by: Selitskiy, Anton, et al.
Published: (2025)
by: Selitskiy, Anton, et al.
Published: (2025)
Improving Generalization for AI-Synthesized Voice Detection
by: Ren, Hainan, et al.
Published: (2024)
by: Ren, Hainan, et al.
Published: (2024)
BiSinger: Bilingual Singing Voice Synthesis
by: Zhou, Huali, et al.
Published: (2023)
by: Zhou, Huali, et al.
Published: (2023)
Optimal Transport Maps are Good Voice Converters
by: Asadulaev, Arip, et al.
Published: (2024)
by: Asadulaev, Arip, et al.
Published: (2024)
Cluster and Separate: a GNN Approach to Voice and Staff Prediction for Score Engraving
by: Foscarin, Francesco, et al.
Published: (2024)
by: Foscarin, Francesco, et al.
Published: (2024)
VQalAttent: a Transparent Speech Generation Pipeline based on Transformer-learned VQ-VAE Latent Space
by: Rodriguez, Armani, et al.
Published: (2024)
by: Rodriguez, Armani, et al.
Published: (2024)
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
by: Cui, Jianwei, et al.
Published: (2024)
by: Cui, Jianwei, et al.
Published: (2024)
Bottleneck Transformer-Based Approach for Improved Automatic STOI Score Prediction
by: Amartyaveer, et al.
Published: (2026)
by: Amartyaveer, et al.
Published: (2026)
Wireless Earphone-based Real-Time Monitoring of Breathing Exercises: A Deep Learning Approach
by: Wazir, Hassam Khan, et al.
Published: (2024)
by: Wazir, Hassam Khan, et al.
Published: (2024)
Tessellated Linear Model for Age Prediction from Voice
by: Alharthi, Dareen, et al.
Published: (2025)
by: Alharthi, Dareen, et al.
Published: (2025)
Voice Signal Processing for Machine Learning. The Case of Speaker Isolation
by: Ganchev, Radan
Published: (2024)
by: Ganchev, Radan
Published: (2024)
On the Generation and Removal of Speaker Adversarial Perturbation for Voice-Privacy Protection
by: Guo, Chenyang, et al.
Published: (2024)
by: Guo, Chenyang, et al.
Published: (2024)
StreamVC: Real-Time Low-Latency Voice Conversion
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
by: Janiczek, John, et al.
Published: (2024)
by: Janiczek, John, et al.
Published: (2024)
Similar Items
-
A Concept-based approach to Voice Disorder Detection
by: Ghia, Davide, et al.
Published: (2025) -
MVP: Multi-source Voice Pathology detection
by: Koudounas, Alkis, et al.
Published: (2025) -
A Contrastive Learning Approach to Mitigate Bias in Speech Models
by: Koudounas, Alkis, et al.
Published: (2024) -
"KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
by: Koudounas, Alkis, et al.
Published: (2025) -
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models
by: Koudounas, Alkis, et al.
Published: (2024)