Discrete Optimal Transport and Voice Conversion
Fuente:
arXiv
Saved in:
| Main Authors: | Selitskiy, Anton, Kocharekar, Maitreya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
by: Lee, Philip H., et al.
Published: (2024)
by: Lee, Philip H., et al.
Published: (2024)
Optimal Transport Maps are Good Voice Converters
by: Asadulaev, Arip, et al.
Published: (2024)
by: Asadulaev, Arip, et al.
Published: (2024)
Zero-shot Voice Conversion with Diffusion Transformers
by: Liu, Songting
Published: (2024)
by: Liu, Songting
Published: (2024)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
by: Du, Zongyang, et al.
Published: (2025)
by: Du, Zongyang, et al.
Published: (2025)
Training-Free Voice Conversion with Factorized Optimal Transport
by: Lobashev, Alexander, et al.
Published: (2025)
by: Lobashev, Alexander, et al.
Published: (2025)
StreamVC: Real-Time Low-Latency Voice Conversion
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
by: Suh, Soobin, et al.
Published: (2025)
by: Suh, Soobin, et al.
Published: (2025)
Phoneme Hallucinator: One-shot Voice Conversion via Set Expansion
by: Shan, Siyuan, et al.
Published: (2023)
by: Shan, Siyuan, et al.
Published: (2023)
Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
by: Morrone, Giovanni, et al.
Published: (2023)
by: Morrone, Giovanni, et al.
Published: (2023)
OpenVoice: Versatile Instant Voice Cloning
by: Qin, Zengyi, et al.
Published: (2023)
by: Qin, Zengyi, et al.
Published: (2023)
A Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal Transport
by: Kaloga, Yacouba, et al.
Published: (2025)
by: Kaloga, Yacouba, et al.
Published: (2025)
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
by: Kaneko, Takuhiro, et al.
Published: (2025)
by: Kaneko, Takuhiro, et al.
Published: (2025)
VANPY: Voice Analysis Framework
by: Koushnir, Gregory, et al.
Published: (2025)
by: Koushnir, Gregory, et al.
Published: (2025)
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
by: Kaneko, Takuhiro, et al.
Published: (2026)
by: Kaneko, Takuhiro, et al.
Published: (2026)
Compact Neural TTS Voices for Accessibility
by: Jain, Kunal, et al.
Published: (2025)
by: Jain, Kunal, et al.
Published: (2025)
Speech to Speech Synthesis for Voice Impersonation
by: Johnson, Bjorn, et al.
Published: (2026)
by: Johnson, Bjorn, et al.
Published: (2026)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
Improving Generalization for AI-Synthesized Voice Detection
by: Ren, Hainan, et al.
Published: (2024)
by: Ren, Hainan, et al.
Published: (2024)
BiSinger: Bilingual Singing Voice Synthesis
by: Zhou, Huali, et al.
Published: (2023)
by: Zhou, Huali, et al.
Published: (2023)
A Concept-based approach to Voice Disorder Detection
by: Ghia, Davide, et al.
Published: (2025)
by: Ghia, Davide, et al.
Published: (2025)
Tessellated Linear Model for Age Prediction from Voice
by: Alharthi, Dareen, et al.
Published: (2025)
by: Alharthi, Dareen, et al.
Published: (2025)
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
by: Kaneko, Takuhiro, et al.
Published: (2024)
by: Kaneko, Takuhiro, et al.
Published: (2024)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
Voice Signal Processing for Machine Learning. The Case of Speaker Isolation
by: Ganchev, Radan
Published: (2024)
by: Ganchev, Radan
Published: (2024)
On the Generation and Removal of Speaker Adversarial Perturbation for Voice-Privacy Protection
by: Guo, Chenyang, et al.
Published: (2024)
by: Guo, Chenyang, et al.
Published: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
by: Janiczek, John, et al.
Published: (2024)
by: Janiczek, John, et al.
Published: (2024)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
by: Hajal, Karl El, et al.
Published: (2025)
by: Hajal, Karl El, et al.
Published: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
by: Hajal, Karl El, et al.
Published: (2025)
by: Hajal, Karl El, et al.
Published: (2025)
Towards General-Purpose Text-Instruction-Guided Voice Conversion
by: Kuan, Chun-Yi, et al.
Published: (2023)
by: Kuan, Chun-Yi, et al.
Published: (2023)
On-device Streaming Discrete Speech Units
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
CoMoSVC: Consistency Model-based Singing Voice Conversion
by: Lu, Yiwen, et al.
Published: (2024)
by: Lu, Yiwen, et al.
Published: (2024)
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
by: Narain, Jaya, et al.
Published: (2025)
by: Narain, Jaya, et al.
Published: (2025)
Evaluating Echo State Network for Parkinson's Disease Prediction using Voice Features
by: Hosseininian, Seyedeh Zahra Seyedi, et al.
Published: (2024)
by: Hosseininian, Seyedeh Zahra Seyedi, et al.
Published: (2024)
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
by: Song, Yulin, et al.
Published: (2024)
by: Song, Yulin, et al.
Published: (2024)
Drax: Speech Recognition with Discrete Flow Matching
by: Navon, Aviv, et al.
Published: (2025)
by: Navon, Aviv, et al.
Published: (2025)
Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
by: Kunze, Tarek, et al.
Published: (2025)
by: Kunze, Tarek, et al.
Published: (2025)
Systematic FAIRness Assessment of Open Voice Biomarker Datasets for Mental Health and Neurodegenerative Diseases
by: Mahapatra, Ishaan, et al.
Published: (2025)
by: Mahapatra, Ishaan, et al.
Published: (2025)
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation
by: Ellinas, Nikolaos, et al.
Published: (2022)
by: Ellinas, Nikolaos, et al.
Published: (2022)
TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
by: Quamer, Waris, et al.
Published: (2026)
by: Quamer, Waris, et al.
Published: (2026)
Similar Items
-
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
by: Lee, Philip H., et al.
Published: (2024) -
Optimal Transport Maps are Good Voice Converters
by: Asadulaev, Arip, et al.
Published: (2024) -
Zero-shot Voice Conversion with Diffusion Transformers
by: Liu, Songting
Published: (2024) -
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
by: Du, Zongyang, et al.
Published: (2025) -
Training-Free Voice Conversion with Factorized Optimal Transport
by: Lobashev, Alexander, et al.
Published: (2025)