MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kaneko, Takuhiro, Kameoka, Hirokazu, Tanaka, Kou, Kondo, Yuto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
Vocoder-Projected Feature Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2020)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2020)
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
Selecting N-lowest scores for training MOS prediction models
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
CoMoSVC: Consistency Model-based Singing Voice Conversion
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
Mitigating Unauthorized Speech Synthesis for Voice Protection
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2024)
Phoneme Hallucinator: One-shot Voice Conversion via Set Expansion
von: Shan, Siyuan, et al.
Veröffentlicht: (2023)
von: Shan, Siyuan, et al.
Veröffentlicht: (2023)
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
von: Nejad, Mahsa Ghazvini, et al.
Veröffentlicht: (2025)
von: Nejad, Mahsa Ghazvini, et al.
Veröffentlicht: (2025)
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
Discrete Optimal Transport and Voice Conversion
von: Selitskiy, Anton, et al.
Veröffentlicht: (2025)
von: Selitskiy, Anton, et al.
Veröffentlicht: (2025)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization
von: Shao, Keren, et al.
Veröffentlicht: (2025)
von: Shao, Keren, et al.
Veröffentlicht: (2025)
When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
von: Kang, Minsu, et al.
Veröffentlicht: (2025)
von: Kang, Minsu, et al.
Veröffentlicht: (2025)
Reproducible Machine Learning-based Voice Pathology Detection: Introducing the Pitch Difference Feature
von: Vrba, Jan, et al.
Veröffentlicht: (2024)
von: Vrba, Jan, et al.
Veröffentlicht: (2024)
Compose Yourself: Average-Velocity Flow Matching for One-Step Speech Enhancement
von: Yang, Gang, et al.
Veröffentlicht: (2025)
von: Yang, Gang, et al.
Veröffentlicht: (2025)
Zero-shot Voice Conversion with Diffusion Transformers
von: Liu, Songting
Veröffentlicht: (2024)
von: Liu, Songting
Veröffentlicht: (2024)
AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
von: Yun, Minhyeok, et al.
Veröffentlicht: (2026)
von: Yun, Minhyeok, et al.
Veröffentlicht: (2026)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2022)
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2022)
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
von: Huang, Yubo, et al.
Veröffentlicht: (2024)
von: Huang, Yubo, et al.
Veröffentlicht: (2024)
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025) -
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024) -
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025) -
Vocoder-Projected Feature Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025) -
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)