Voice Attribute Editing with Text Prompt
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sheng, Zhengyan, Ai, Yang, Liu, Li-Juan, Pan, Jia, Ling, Zhen-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
von: Wang, Rui, et al.
Veröffentlicht: (2024)
von: Wang, Rui, et al.
Veröffentlicht: (2024)
Introducing voice timbre attribute detection
von: He, Jinghao, et al.
Veröffentlicht: (2025)
von: He, Jinghao, et al.
Veröffentlicht: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2024)
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2024)
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
von: Liu, Fei, et al.
Veröffentlicht: (2024)
von: Liu, Fei, et al.
Veröffentlicht: (2024)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
von: Du, Zongcai, et al.
Veröffentlicht: (2025)
von: Du, Zongcai, et al.
Veröffentlicht: (2025)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
A Real-Time Voice Activity Detection Based On Lightweight Neural
von: Jia, Jidong, et al.
Veröffentlicht: (2024)
von: Jia, Jidong, et al.
Veröffentlicht: (2024)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
von: Guragain, Anmol, et al.
Veröffentlicht: (2024)
von: Guragain, Anmol, et al.
Veröffentlicht: (2024)
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2024)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2024)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
Leveraging Prompt Learning and Pause Encoding for Alzheimer's Disease Detection
von: Liu, Yin-Long, et al.
Veröffentlicht: (2024)
von: Liu, Yin-Long, et al.
Veröffentlicht: (2024)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
In-the-wild Audio Spatialization with Flexible Text-guided Localization
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
von: Yang, Yuguang, et al.
Veröffentlicht: (2024)
von: Yang, Yuguang, et al.
Veröffentlicht: (2024)
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2025)
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2025)
Voice Cloning: Comprehensive Survey
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025)
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
Universal Preference-Score-based Pairwise Speech Quality Assessment
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
A New Approach to Voice Authenticity
von: Müller, Nicolas M., et al.
Veröffentlicht: (2024)
von: Müller, Nicolas M., et al.
Veröffentlicht: (2024)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025) -
Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025) -
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
von: Wang, Rui, et al.
Veröffentlicht: (2024) -
Introducing voice timbre attribute detection
von: He, Jinghao, et al.
Veröffentlicht: (2025) -
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)