VANPY: Voice Analysis Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Koushnir, Gregory, Fire, Michael, Alpert, Galit Fuhrmann, Kagan, Dima |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OpenVoice: Versatile Instant Voice Cloning
di: Qin, Zengyi, et al.
Pubblicazione: (2023)
di: Qin, Zengyi, et al.
Pubblicazione: (2023)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
di: Du, Zongyang, et al.
Pubblicazione: (2025)
di: Du, Zongyang, et al.
Pubblicazione: (2025)
Compact Neural TTS Voices for Accessibility
di: Jain, Kunal, et al.
Pubblicazione: (2025)
di: Jain, Kunal, et al.
Pubblicazione: (2025)
Discrete Optimal Transport and Voice Conversion
di: Selitskiy, Anton, et al.
Pubblicazione: (2025)
di: Selitskiy, Anton, et al.
Pubblicazione: (2025)
Speech to Speech Synthesis for Voice Impersonation
di: Johnson, Bjorn, et al.
Pubblicazione: (2026)
di: Johnson, Bjorn, et al.
Pubblicazione: (2026)
Improving Generalization for AI-Synthesized Voice Detection
di: Ren, Hainan, et al.
Pubblicazione: (2024)
di: Ren, Hainan, et al.
Pubblicazione: (2024)
BiSinger: Bilingual Singing Voice Synthesis
di: Zhou, Huali, et al.
Pubblicazione: (2023)
di: Zhou, Huali, et al.
Pubblicazione: (2023)
Zero-shot Voice Conversion with Diffusion Transformers
di: Liu, Songting
Pubblicazione: (2024)
di: Liu, Songting
Pubblicazione: (2024)
Optimal Transport Maps are Good Voice Converters
di: Asadulaev, Arip, et al.
Pubblicazione: (2024)
di: Asadulaev, Arip, et al.
Pubblicazione: (2024)
A Concept-based approach to Voice Disorder Detection
di: Ghia, Davide, et al.
Pubblicazione: (2025)
di: Ghia, Davide, et al.
Pubblicazione: (2025)
Tessellated Linear Model for Age Prediction from Voice
di: Alharthi, Dareen, et al.
Pubblicazione: (2025)
di: Alharthi, Dareen, et al.
Pubblicazione: (2025)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2026)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2026)
Voice Signal Processing for Machine Learning. The Case of Speaker Isolation
di: Ganchev, Radan
Pubblicazione: (2024)
di: Ganchev, Radan
Pubblicazione: (2024)
On the Generation and Removal of Speaker Adversarial Perturbation for Voice-Privacy Protection
di: Guo, Chenyang, et al.
Pubblicazione: (2024)
di: Guo, Chenyang, et al.
Pubblicazione: (2024)
StreamVC: Real-Time Low-Latency Voice Conversion
di: Yang, Yang, et al.
Pubblicazione: (2024)
di: Yang, Yang, et al.
Pubblicazione: (2024)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
di: Lee, Philip H., et al.
Pubblicazione: (2024)
di: Lee, Philip H., et al.
Pubblicazione: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
di: Janiczek, John, et al.
Pubblicazione: (2024)
di: Janiczek, John, et al.
Pubblicazione: (2024)
Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
di: Suh, Soobin, et al.
Pubblicazione: (2025)
di: Suh, Soobin, et al.
Pubblicazione: (2025)
Phoneme Hallucinator: One-shot Voice Conversion via Set Expansion
di: Shan, Siyuan, et al.
Pubblicazione: (2023)
di: Shan, Siyuan, et al.
Pubblicazione: (2023)
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
di: Narain, Jaya, et al.
Pubblicazione: (2025)
di: Narain, Jaya, et al.
Pubblicazione: (2025)
Evaluating Echo State Network for Parkinson's Disease Prediction using Voice Features
di: Hosseininian, Seyedeh Zahra Seyedi, et al.
Pubblicazione: (2024)
di: Hosseininian, Seyedeh Zahra Seyedi, et al.
Pubblicazione: (2024)
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
di: Song, Yulin, et al.
Pubblicazione: (2024)
di: Song, Yulin, et al.
Pubblicazione: (2024)
Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
di: Kunze, Tarek, et al.
Pubblicazione: (2025)
di: Kunze, Tarek, et al.
Pubblicazione: (2025)
Systematic FAIRness Assessment of Open Voice Biomarker Datasets for Mental Health and Neurodegenerative Diseases
di: Mahapatra, Ishaan, et al.
Pubblicazione: (2025)
di: Mahapatra, Ishaan, et al.
Pubblicazione: (2025)
Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features
di: Chen, Wei, et al.
Pubblicazione: (2025)
di: Chen, Wei, et al.
Pubblicazione: (2025)
RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations
di: Kim, Seungmin, et al.
Pubblicazione: (2025)
di: Kim, Seungmin, et al.
Pubblicazione: (2025)
SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
di: Tang, Beilong, et al.
Pubblicazione: (2025)
di: Tang, Beilong, et al.
Pubblicazione: (2025)
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
di: Yin, Chun, et al.
Pubblicazione: (2024)
di: Yin, Chun, et al.
Pubblicazione: (2024)
Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
di: Ariyanti, Whenty, et al.
Pubblicazione: (2025)
di: Ariyanti, Whenty, et al.
Pubblicazione: (2025)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
di: Morrone, Giovanni, et al.
Pubblicazione: (2023)
di: Morrone, Giovanni, et al.
Pubblicazione: (2023)
Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations
di: Kakoulidis, Panos, et al.
Pubblicazione: (2024)
di: Kakoulidis, Panos, et al.
Pubblicazione: (2024)
CogniVoice: Multimodal and Multilingual Fusion Networks for Mild Cognitive Impairment Assessment from Spontaneous Speech
di: Cheng, Jiali, et al.
Pubblicazione: (2024)
di: Cheng, Jiali, et al.
Pubblicazione: (2024)
Voice-Driven Mortality Prediction in Hospitalized Heart Failure Patients: A Machine Learning Approach Enhanced with Diagnostic Biomarkers
di: Ahmadli, Nihat, et al.
Pubblicazione: (2024)
di: Ahmadli, Nihat, et al.
Pubblicazione: (2024)
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
di: Cui, Jianwei, et al.
Pubblicazione: (2024)
di: Cui, Jianwei, et al.
Pubblicazione: (2024)
Prosody Analysis of Audiobooks
di: Pethe, Charuta, et al.
Pubblicazione: (2023)
di: Pethe, Charuta, et al.
Pubblicazione: (2023)
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2026)
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2026)
Advanced Framework for Animal Sound Classification With Features Optimization
di: Yang, Qiang, et al.
Pubblicazione: (2024)
di: Yang, Qiang, et al.
Pubblicazione: (2024)
Self-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling
di: Lepage, Theo, et al.
Pubblicazione: (2025)
di: Lepage, Theo, et al.
Pubblicazione: (2025)
A Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal Transport
di: Kaloga, Yacouba, et al.
Pubblicazione: (2025)
di: Kaloga, Yacouba, et al.
Pubblicazione: (2025)
Additive Margin in Contrastive Self-Supervised Frameworks to Learn Discriminative Speaker Representations
di: Lepage, Theo, et al.
Pubblicazione: (2024)
di: Lepage, Theo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
OpenVoice: Versatile Instant Voice Cloning
di: Qin, Zengyi, et al.
Pubblicazione: (2023) -
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
di: Du, Zongyang, et al.
Pubblicazione: (2025) -
Compact Neural TTS Voices for Accessibility
di: Jain, Kunal, et al.
Pubblicazione: (2025) -
Discrete Optimal Transport and Voice Conversion
di: Selitskiy, Anton, et al.
Pubblicazione: (2025) -
Speech to Speech Synthesis for Voice Impersonation
di: Johnson, Bjorn, et al.
Pubblicazione: (2026)