EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Ziqi, Wang, Jianzong, Zhang, Xulong, Zhang, Yong, Cheng, Ning, Xiao, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN Optimization
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation
von: Liang, Ziqi, et al.
Veröffentlicht: (2025)
von: Liang, Ziqi, et al.
Veröffentlicht: (2025)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
von: Wang, Jianzong, et al.
Veröffentlicht: (2023)
von: Wang, Jianzong, et al.
Veröffentlicht: (2023)
RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion
von: Sim, Youngjun, et al.
Veröffentlicht: (2024)
von: Sim, Youngjun, et al.
Veröffentlicht: (2024)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
von: Yang, Qingran, et al.
Veröffentlicht: (2026)
von: Yang, Qingran, et al.
Veröffentlicht: (2026)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
von: Li, Yuke, et al.
Veröffentlicht: (2024)
von: Li, Yuke, et al.
Veröffentlicht: (2024)
EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
von: Zhang, Bingyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Bingyuan, et al.
Veröffentlicht: (2024)
ESARM: 3D Emotional Speech-to-Animation via Reward Model from Automatically-Ranked Demonstrations
von: Zhang, Xulong, et al.
Veröffentlicht: (2024)
von: Zhang, Xulong, et al.
Veröffentlicht: (2024)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
von: Tu, Huu Tuong, et al.
Veröffentlicht: (2025)
von: Tu, Huu Tuong, et al.
Veröffentlicht: (2025)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
AdaptVC: High Quality Voice Conversion with Adaptive Learning
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
StreamVC: Real-Time Low-Latency Voice Conversion
von: Yang, Yang, et al.
Veröffentlicht: (2024)
von: Yang, Yang, et al.
Veröffentlicht: (2024)
RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
von: Bargum, Anders R., et al.
Veröffentlicht: (2024)
von: Bargum, Anders R., et al.
Veröffentlicht: (2024)
Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning
von: Yang, Qian, et al.
Veröffentlicht: (2025)
von: Yang, Qian, et al.
Veröffentlicht: (2025)
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
von: Yang, Yuguang, et al.
Veröffentlicht: (2024)
von: Yang, Yuguang, et al.
Veröffentlicht: (2024)
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
SelfVC: Voice Conversion With Iterative Refinement using Self Transformations
von: Neekhara, Paarth, et al.
Veröffentlicht: (2023)
von: Neekhara, Paarth, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
von: Li, Pengcheng, et al.
Veröffentlicht: (2024) -
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
von: Deng, Yimin, et al.
Veröffentlicht: (2024) -
EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN Optimization
von: Wang, Jianzong, et al.
Veröffentlicht: (2024) -
Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
von: Deng, Yimin, et al.
Veröffentlicht: (2024) -
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)