GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Fan, Zhao, Kaicheng, Fleisch, Elgar, Barata, Filipe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Patient-Level Multimodal Question Answering from Multi-Site Auscultation Recordings
by: Wu, Fan, et al.
Published: (2026)
by: Wu, Fan, et al.
Published: (2026)
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
by: Togootogtokh, Enkhtogtokh, et al.
Published: (2025)
by: Togootogtokh, Enkhtogtokh, et al.
Published: (2025)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
by: Sorrenti, Adam
Published: (2024)
by: Sorrenti, Adam
Published: (2024)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
by: Pan, Yu, et al.
Published: (2024)
by: Pan, Yu, et al.
Published: (2024)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
by: Yao, Wenhan, et al.
Published: (2024)
by: Yao, Wenhan, et al.
Published: (2024)
A Real-Time Voice Activity Detection Based On Lightweight Neural
by: Jia, Jidong, et al.
Published: (2024)
by: Jia, Jidong, et al.
Published: (2024)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
by: Xu, Rixi, et al.
Published: (2026)
by: Xu, Rixi, et al.
Published: (2026)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
by: Kim, Nam-Gyu, et al.
Published: (2025)
by: Kim, Nam-Gyu, et al.
Published: (2025)
LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning
by: Dhar, Sandipan, et al.
Published: (2025)
by: Dhar, Sandipan, et al.
Published: (2025)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
by: Zhao, Junchuan, et al.
Published: (2025)
by: Zhao, Junchuan, et al.
Published: (2025)
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
by: Huang, Yubo, et al.
Published: (2024)
by: Huang, Yubo, et al.
Published: (2024)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
by: Lau, Hok-Shing, et al.
Published: (2024)
by: Lau, Hok-Shing, et al.
Published: (2024)
Voice Cloning: Comprehensive Survey
by: Azzuni, Hussam, et al.
Published: (2025)
by: Azzuni, Hussam, et al.
Published: (2025)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
by: Huang, Jiawei, et al.
Published: (2024)
by: Huang, Jiawei, et al.
Published: (2024)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
by: Qian, Zhiwen, et al.
Published: (2025)
by: Qian, Zhiwen, et al.
Published: (2025)
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
by: Yang, Yuguang, et al.
Published: (2024)
by: Yang, Yuguang, et al.
Published: (2024)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
by: Anastassiou, Philip, et al.
Published: (2024)
by: Anastassiou, Philip, et al.
Published: (2024)
Voice Attribute Editing with Text Prompt
by: Sheng, Zhengyan, et al.
Published: (2024)
by: Sheng, Zhengyan, et al.
Published: (2024)
A New Approach to Voice Authenticity
by: Müller, Nicolas M., et al.
Published: (2024)
by: Müller, Nicolas M., et al.
Published: (2024)
Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models
by: Nikolayevich, Borodin Kirill, et al.
Published: (2024)
by: Nikolayevich, Borodin Kirill, et al.
Published: (2024)
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
by: An, Keyu, et al.
Published: (2024)
by: An, Keyu, et al.
Published: (2024)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
by: Du, Zhihao, et al.
Published: (2025)
by: Du, Zhihao, et al.
Published: (2025)
Anonymising Elderly and Pathological Speech: Voice Conversion Using DDSP and Query-by-Example
by: Ghosh, Suhita, et al.
Published: (2024)
by: Ghosh, Suhita, et al.
Published: (2024)
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
by: Qian, Jiale, et al.
Published: (2026)
by: Qian, Jiale, et al.
Published: (2026)
SingFake: Singing Voice Deepfake Detection
by: Zang, Yongyi, et al.
Published: (2023)
by: Zang, Yongyi, et al.
Published: (2023)
Deepfake Detection of Singing Voices With Whisper Encodings
by: Sharma, Falguni, et al.
Published: (2025)
by: Sharma, Falguni, et al.
Published: (2025)
Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation
by: Ji, Shulei, et al.
Published: (2025)
by: Ji, Shulei, et al.
Published: (2025)
LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
by: Zou, Wenhao, et al.
Published: (2026)
by: Zou, Wenhao, et al.
Published: (2026)
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
by: Wang, Kaidi, et al.
Published: (2025)
by: Wang, Kaidi, et al.
Published: (2025)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
by: Guragain, Anmol, et al.
Published: (2024)
by: Guragain, Anmol, et al.
Published: (2024)
Physics-Guided Deepfake Detection for Voice Authentication Systems
by: Mohammadi, Alireza, et al.
Published: (2025)
by: Mohammadi, Alireza, et al.
Published: (2025)
Reproducible Machine Learning-based Voice Pathology Detection: Introducing the Pitch Difference Feature
by: Vrba, Jan, et al.
Published: (2024)
by: Vrba, Jan, et al.
Published: (2024)
Developing a Multi-variate Prediction Model For COVID-19 From Crowd-sourced Respiratory Voice Data
by: Yan, Yuyang, et al.
Published: (2024)
by: Yan, Yuyang, et al.
Published: (2024)
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
by: Valdivia, Andrew, et al.
Published: (2025)
by: Valdivia, Andrew, et al.
Published: (2025)
The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
by: Sheng, Zhengyan, et al.
Published: (2025)
by: Sheng, Zhengyan, et al.
Published: (2025)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
by: Wang, Rui, et al.
Published: (2024)
by: Wang, Rui, et al.
Published: (2024)
RDSinger: Reference-based Diffusion Network for Singing Voice Synthesis
by: Sui, Kehan, et al.
Published: (2024)
by: Sui, Kehan, et al.
Published: (2024)
Similar Items
-
Patient-Level Multimodal Question Answering from Multi-Site Auscultation Recordings
by: Wu, Fan, et al.
Published: (2026) -
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
by: Togootogtokh, Enkhtogtokh, et al.
Published: (2025) -
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
by: Moell, Birger, et al.
Published: (2025) -
Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
by: Sorrenti, Adam
Published: (2024) -
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
by: Pan, Yu, et al.
Published: (2024)