Fed-PISA: Federated Voice Cloning via Personalized Identity-Style Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Qi, Ma, Shituo, Yu, Guoxin, Peng, Hanyang, Yu, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
Voice Cloning: Comprehensive Survey
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025)
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
von: Valdivia, Andrew, et al.
Veröffentlicht: (2025)
von: Valdivia, Andrew, et al.
Veröffentlicht: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
von: Moell, Birger, et al.
Veröffentlicht: (2025)
von: Moell, Birger, et al.
Veröffentlicht: (2025)
VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
von: Zhan, Jun, et al.
Veröffentlicht: (2025)
von: Zhan, Jun, et al.
Veröffentlicht: (2025)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
von: Zhou, Fangru, et al.
Veröffentlicht: (2025)
von: Zhou, Fangru, et al.
Veröffentlicht: (2025)
Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
von: Zhou, Xukun, et al.
Veröffentlicht: (2024)
von: Zhou, Xukun, et al.
Veröffentlicht: (2024)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
von: Du, Zhihao, et al.
Veröffentlicht: (2025)
von: Du, Zhihao, et al.
Veröffentlicht: (2025)
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
von: Yang, Yuguang, et al.
Veröffentlicht: (2024)
von: Yang, Yuguang, et al.
Veröffentlicht: (2024)
LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning
von: Dhar, Sandipan, et al.
Veröffentlicht: (2025)
von: Dhar, Sandipan, et al.
Veröffentlicht: (2025)
USM-VC: Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block for Voice Conversion
von: Li, Na, et al.
Veröffentlicht: (2025)
von: Li, Na, et al.
Veröffentlicht: (2025)
LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
von: Zou, Wenhao, et al.
Veröffentlicht: (2026)
von: Zou, Wenhao, et al.
Veröffentlicht: (2026)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction
von: Li, Yadong, et al.
Veröffentlicht: (2026)
von: Li, Yadong, et al.
Veröffentlicht: (2026)
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
von: Chen, Yun, et al.
Veröffentlicht: (2023)
von: Chen, Yun, et al.
Veröffentlicht: (2023)
OpenVoice: Versatile Instant Voice Cloning
von: Qin, Zengyi, et al.
Veröffentlicht: (2023)
von: Qin, Zengyi, et al.
Veröffentlicht: (2023)
De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks
von: Fan, Wei, et al.
Veröffentlicht: (2025)
von: Fan, Wei, et al.
Veröffentlicht: (2025)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
von: Qi, Xin, et al.
Veröffentlicht: (2024)
von: Qi, Xin, et al.
Veröffentlicht: (2024)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
von: Li, Haitao, et al.
Veröffentlicht: (2026)
von: Li, Haitao, et al.
Veröffentlicht: (2026)
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
von: Nejad, Mahsa Ghazvini, et al.
Veröffentlicht: (2025)
von: Nejad, Mahsa Ghazvini, et al.
Veröffentlicht: (2025)
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
von: Meng, Ming, et al.
Veröffentlicht: (2025)
von: Meng, Ming, et al.
Veröffentlicht: (2025)
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2025)
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2025)
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
BNMusic: Blending Environmental Noises into Personalized Music
von: Zuo, Chi, et al.
Veröffentlicht: (2025)
von: Zuo, Chi, et al.
Veröffentlicht: (2025)
A Real-Time Voice Activity Detection Based On Lightweight Neural
von: Jia, Jidong, et al.
Veröffentlicht: (2024)
von: Jia, Jidong, et al.
Veröffentlicht: (2024)
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
von: Yun, Minhyeok, et al.
Veröffentlicht: (2026)
von: Yun, Minhyeok, et al.
Veröffentlicht: (2026)
DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
von: Du, Zongcai, et al.
Veröffentlicht: (2025)
von: Du, Zongcai, et al.
Veröffentlicht: (2025)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
von: Wang, Rui, et al.
Veröffentlicht: (2024)
von: Wang, Rui, et al.
Veröffentlicht: (2024)
Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
von: Li, Guojian, et al.
Veröffentlicht: (2025)
von: Li, Guojian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
von: Li, Haiyun, et al.
Veröffentlicht: (2025) -
Voice Cloning: Comprehensive Survey
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025) -
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
von: Xu, Rixi, et al.
Veröffentlicht: (2026) -
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
von: Li, Renyuan, et al.
Veröffentlicht: (2025) -
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)