VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kameoka, Hirokazu, Kaneko, Takuhiro, Tanaka, Kou, Hojo, Nobukatsu, Seki, Shogo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2020
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2026)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2026)
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Selecting N-lowest scores for training MOS prediction models
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Vocoder-Projected Feature Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
von: Tu, Huu Tuong, et al.
Veröffentlicht: (2025)
von: Tu, Huu Tuong, et al.
Veröffentlicht: (2025)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2026)
von: Wang, Zhichao, et al.
Veröffentlicht: (2026)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
Improved Remixing Process for Domain Adaptation-Based Speech Enhancement by Mitigating Data Imbalance in Signal-to-Noise Ratio
von: Li, Li, et al.
Veröffentlicht: (2024)
von: Li, Li, et al.
Veröffentlicht: (2024)
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
von: Kim, Heeseung, et al.
Veröffentlicht: (2025)
von: Kim, Heeseung, et al.
Veröffentlicht: (2025)
Residual Speaker Representation for One-Shot Voice Conversion
von: Xu, Le, et al.
Veröffentlicht: (2023)
von: Xu, Le, et al.
Veröffentlicht: (2023)
Voice Conversion-based Privacy through Adversarial Information Hiding
von: Webber, Jacob J, et al.
Veröffentlicht: (2024)
von: Webber, Jacob J, et al.
Veröffentlicht: (2024)
RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
von: Bargum, Anders R., et al.
Veröffentlicht: (2024)
von: Bargum, Anders R., et al.
Veröffentlicht: (2024)
Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2026)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2026)
The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
An Extensive Analysis of the Singing Voice Conversion Challenge 2025 Evaluation Results
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
von: Xu, Haiying, et al.
Veröffentlicht: (2025)
von: Xu, Haiying, et al.
Veröffentlicht: (2025)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025) -
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025) -
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024) -
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2026) -
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)