Enregistré dans:
| Auteurs principaux: | Chen, Wuyang, Sun, Yanjie, Xu, Kele, Dou, Yong |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2408.02025 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
par: Deng, Xuyao, et autres
Publié: (2025)
par: Deng, Xuyao, et autres
Publié: (2025)
VoiceWukong: Benchmarking Deepfake Voice Detection
par: Yan, Ziwei, et autres
Publié: (2024)
par: Yan, Ziwei, et autres
Publié: (2024)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
par: Du, Zhihao, et autres
Publié: (2024)
par: Du, Zhihao, et autres
Publié: (2024)
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
par: Park, Kyeongman, et autres
Publié: (2025)
par: Park, Kyeongman, et autres
Publié: (2025)
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
par: Chen, Yun, et autres
Publié: (2023)
par: Chen, Yun, et autres
Publié: (2023)
Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
par: Dai, Weinan, et autres
Publié: (2024)
par: Dai, Weinan, et autres
Publié: (2024)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
par: Xu, Rixi, et autres
Publié: (2026)
par: Xu, Rixi, et autres
Publié: (2026)
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
par: Yun, Minhyeok, et autres
Publié: (2026)
par: Yun, Minhyeok, et autres
Publié: (2026)
Towards Attention-based Contrastive Learning for Audio Spoof Detection
par: Goel, Chirag, et autres
Publié: (2024)
par: Goel, Chirag, et autres
Publié: (2024)
LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
par: Zou, Wenhao, et autres
Publié: (2026)
par: Zou, Wenhao, et autres
Publié: (2026)
Collective Learning Mechanism based Optimal Transport Generative Adversarial Network for Non-parallel Voice Conversion
par: Dhar, Sandipan, et autres
Publié: (2025)
par: Dhar, Sandipan, et autres
Publié: (2025)
Safe Guard: an LLM-agent for Real-time Voice-based Hate Speech Detection in Social Virtual Reality
par: Xu, Yiwen, et autres
Publié: (2024)
par: Xu, Yiwen, et autres
Publié: (2024)
RDSinger: Reference-based Diffusion Network for Singing Voice Synthesis
par: Sui, Kehan, et autres
Publié: (2024)
par: Sui, Kehan, et autres
Publié: (2024)
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
par: Wang, Kaidi, et autres
Publié: (2025)
par: Wang, Kaidi, et autres
Publié: (2025)
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
par: Valdivia, Andrew, et autres
Publié: (2025)
par: Valdivia, Andrew, et autres
Publié: (2025)
Generative Adversarial Network based Voice Conversion: Techniques, Challenges, and Recent Advancements
par: Dhar, Sandipan, et autres
Publié: (2025)
par: Dhar, Sandipan, et autres
Publié: (2025)
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
par: Togootogtokh, Enkhtogtokh, et autres
Publié: (2025)
par: Togootogtokh, Enkhtogtokh, et autres
Publié: (2025)
Voice Cloning: Comprehensive Survey
par: Azzuni, Hussam, et autres
Publié: (2025)
par: Azzuni, Hussam, et autres
Publié: (2025)
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
par: Su, Yi, et autres
Publié: (2025)
par: Su, Yi, et autres
Publié: (2025)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
par: Wang, Rui, et autres
Publié: (2024)
par: Wang, Rui, et autres
Publié: (2024)
The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
par: Sheng, Zhengyan, et autres
Publié: (2025)
par: Sheng, Zhengyan, et autres
Publié: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
par: Anastassiou, Philip, et autres
Publié: (2024)
par: Anastassiou, Philip, et autres
Publié: (2024)
DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice
par: Zhang, Leying, et autres
Publié: (2026)
par: Zhang, Leying, et autres
Publié: (2026)
SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
par: Qian, Jiale, et autres
Publié: (2026)
par: Qian, Jiale, et autres
Publié: (2026)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
par: Zhao, Junchuan, et autres
Publié: (2025)
par: Zhao, Junchuan, et autres
Publié: (2025)
Voice Attribute Editing with Text Prompt
par: Sheng, Zhengyan, et autres
Publié: (2024)
par: Sheng, Zhengyan, et autres
Publié: (2024)
A New Approach to Voice Authenticity
par: Müller, Nicolas M., et autres
Publié: (2024)
par: Müller, Nicolas M., et autres
Publié: (2024)
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
par: Huang, Yubo, et autres
Publié: (2024)
par: Huang, Yubo, et autres
Publié: (2024)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
par: Huang, Jiawei, et autres
Publié: (2024)
par: Huang, Jiawei, et autres
Publié: (2024)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
par: Zhang, Chi, et autres
Publié: (2025)
par: Zhang, Chi, et autres
Publié: (2025)
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
par: Watcharasupat, Karn N., et autres
Publié: (2024)
par: Watcharasupat, Karn N., et autres
Publié: (2024)
R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
par: Zheng, Junjie, et autres
Publié: (2025)
par: Zheng, Junjie, et autres
Publié: (2025)
SingFake: Singing Voice Deepfake Detection
par: Zang, Yongyi, et autres
Publié: (2023)
par: Zang, Yongyi, et autres
Publié: (2023)
Deepfake Detection of Singing Voices With Whisper Encodings
par: Sharma, Falguni, et autres
Publié: (2025)
par: Sharma, Falguni, et autres
Publié: (2025)
Contrastive Learning with Spectrum Information Augmentation in Abnormal Sound Detection
par: Meng, Xinxin, et autres
Publié: (2025)
par: Meng, Xinxin, et autres
Publié: (2025)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
par: Choi, Ha-Yeong, et autres
Publié: (2025)
par: Choi, Ha-Yeong, et autres
Publié: (2025)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
par: Du, Zhihao, et autres
Publié: (2025)
par: Du, Zhihao, et autres
Publié: (2025)
Prior-agnostic Multi-scale Contrastive Text-Audio Pre-training for Parallelized TTS Frontend Modeling
par: Wang, Quanxiu, et autres
Publié: (2024)
par: Wang, Quanxiu, et autres
Publié: (2024)
Physics-Guided Deepfake Detection for Voice Authentication Systems
par: Mohammadi, Alireza, et autres
Publié: (2025)
par: Mohammadi, Alireza, et autres
Publié: (2025)
AudioCIL: A Python Toolbox for Audio Class-Incremental Learning with Multiple Scenes
par: Xu, Qisheng, et autres
Publié: (2024)
par: Xu, Qisheng, et autres
Publié: (2024)
Documents similaires
-
Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
par: Deng, Xuyao, et autres
Publié: (2025) -
VoiceWukong: Benchmarking Deepfake Voice Detection
par: Yan, Ziwei, et autres
Publié: (2024) -
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
par: Du, Zhihao, et autres
Publié: (2024) -
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
par: Park, Kyeongman, et autres
Publié: (2025) -
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
par: Chen, Yun, et autres
Publié: (2023)