SelfVC: Voice Conversion With Iterative Refinement using Self Transformations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Neekhara, Paarth, Hussain, Shehzeen, Valle, Rafael, Ginsburg, Boris, Ranjan, Rishabh, Dubnov, Shlomo, Koushanfar, Farinaz, McAuley, Julian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
von: Neekhara, Paarth, et al.
Veröffentlicht: (2024)
von: Neekhara, Paarth, et al.
Veröffentlicht: (2024)
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
von: Langman, Ryan, et al.
Veröffentlicht: (2025)
von: Langman, Ryan, et al.
Veröffentlicht: (2025)
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2025)
von: Casanova, Edresson, et al.
Veröffentlicht: (2025)
Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
von: Xu, Weihan, et al.
Veröffentlicht: (2024)
von: Xu, Weihan, et al.
Veröffentlicht: (2024)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
von: Hussain, Shehzeen, et al.
Veröffentlicht: (2025)
von: Hussain, Shehzeen, et al.
Veröffentlicht: (2025)
Improving Music Source Separation with Diffusion and Consistency Refinement
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
GenVC: Self-Supervised Zero-Shot Voice Conversion
von: Cai, Zexin, et al.
Veröffentlicht: (2025)
von: Cai, Zexin, et al.
Veröffentlicht: (2025)
kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization
von: Shao, Keren, et al.
Veröffentlicht: (2025)
von: Shao, Keren, et al.
Veröffentlicht: (2025)
REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models
von: Zhang, Ruisi, et al.
Veröffentlicht: (2023)
von: Zhang, Ruisi, et al.
Veröffentlicht: (2023)
Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation
von: Fejgin, Roy, et al.
Veröffentlicht: (2025)
von: Fejgin, Roy, et al.
Veröffentlicht: (2025)
Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
von: Hussain, Shehzeen, et al.
Veröffentlicht: (2025)
von: Hussain, Shehzeen, et al.
Veröffentlicht: (2025)
Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion Models
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
Binaural sound source localization using a hybrid time and frequency domain model
von: Geva, Gil, et al.
Veröffentlicht: (2024)
von: Geva, Gil, et al.
Veröffentlicht: (2024)
BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music
von: Yao, Mingyang, et al.
Veröffentlicht: (2025)
von: Yao, Mingyang, et al.
Veröffentlicht: (2025)
PosCUDA: Position based Convolution for Unlearnable Audio Datasets
von: Gokul, Vignesh, et al.
Veröffentlicht: (2024)
von: Gokul, Vignesh, et al.
Veröffentlicht: (2024)
Interpreting Graphic Notation with MusicLDM: An AI Improvisation of Cornelius Cardew's Treatise
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2025)
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
Music Enhancement with Deep Filters: A Technical Report for The ICASSP 2024 Cadenza Challenge
von: Shao, Keren, et al.
Veröffentlicht: (2024)
von: Shao, Keren, et al.
Veröffentlicht: (2024)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
von: Zheng, Qixi, et al.
Veröffentlicht: (2026)
von: Zheng, Qixi, et al.
Veröffentlicht: (2026)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
AdaptVC: High Quality Voice Conversion with Adaptive Learning
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
StreamVC: Real-Time Low-Latency Voice Conversion
von: Yang, Yang, et al.
Veröffentlicht: (2024)
von: Yang, Yang, et al.
Veröffentlicht: (2024)
DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
von: Li, Yuke, et al.
Veröffentlicht: (2024)
von: Li, Yuke, et al.
Veröffentlicht: (2024)
Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
von: Tu, Huu Tuong, et al.
Veröffentlicht: (2025)
von: Tu, Huu Tuong, et al.
Veröffentlicht: (2025)
DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion
von: Ning, Ziqian, et al.
Veröffentlicht: (2024)
von: Ning, Ziqian, et al.
Veröffentlicht: (2024)
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
von: Neekhara, Paarth, et al.
Veröffentlicht: (2024) -
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
von: Langman, Ryan, et al.
Veröffentlicht: (2025) -
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2025) -
Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
von: Xu, Weihan, et al.
Veröffentlicht: (2024) -
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)