DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
Fuente:
arXiv
Salvato in:
| Autori principali: | Ulgen, Ismail Rasim, Cai, Zexin, Andrews, Nicholas, Koehn, Philipp, Sisman, Berrak |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2025)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2025)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025)
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
di: Lee, Philip H., et al.
Pubblicazione: (2024)
di: Lee, Philip H., et al.
Pubblicazione: (2024)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
di: Du, Zongyang, et al.
Pubblicazione: (2025)
di: Du, Zongyang, et al.
Pubblicazione: (2025)
Can Emotion Fool Anti-spoofing?
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025)
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025)
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
di: Zhao, Xiutian, et al.
Pubblicazione: (2026)
di: Zhao, Xiutian, et al.
Pubblicazione: (2026)
ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2026)
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2026)
Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
di: Du, Zongyang, et al.
Pubblicazione: (2024)
di: Du, Zongyang, et al.
Pubblicazione: (2024)
Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
di: Salman, Ali N., et al.
Pubblicazione: (2024)
di: Salman, Ali N., et al.
Pubblicazione: (2024)
Universal Speech Content Factorization
di: Xinyuan, Henry Li, et al.
Pubblicazione: (2026)
di: Xinyuan, Henry Li, et al.
Pubblicazione: (2026)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2023)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2023)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
PRESENT: Zero-Shot Text-to-Prosody Control
di: Lam, Perry, et al.
Pubblicazione: (2024)
di: Lam, Perry, et al.
Pubblicazione: (2024)
emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2024)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2024)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
di: Lam, Perry, et al.
Pubblicazione: (2022)
di: Lam, Perry, et al.
Pubblicazione: (2022)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
di: Jawaid, Ahad, et al.
Pubblicazione: (2024)
di: Jawaid, Ahad, et al.
Pubblicazione: (2024)
LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning
di: Dhar, Sandipan, et al.
Pubblicazione: (2025)
di: Dhar, Sandipan, et al.
Pubblicazione: (2025)
Versatile audio-visual learning for emotion recognition
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Prosody Analysis of Audiobooks
di: Pethe, Charuta, et al.
Pubblicazione: (2023)
di: Pethe, Charuta, et al.
Pubblicazione: (2023)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
di: Li, Yuke, et al.
Pubblicazione: (2024)
di: Li, Yuke, et al.
Pubblicazione: (2024)
SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
di: Tang, Beilong, et al.
Pubblicazione: (2025)
di: Tang, Beilong, et al.
Pubblicazione: (2025)
Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference
di: Zhang, Yanzhe, et al.
Pubblicazione: (2024)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2024)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
di: Qi, Tianhua, et al.
Pubblicazione: (2024)
di: Qi, Tianhua, et al.
Pubblicazione: (2024)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
di: Xu, Rixi, et al.
Pubblicazione: (2026)
di: Xu, Rixi, et al.
Pubblicazione: (2026)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
di: Zhang, Lin, et al.
Pubblicazione: (2026)
di: Zhang, Lin, et al.
Pubblicazione: (2026)
The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization
di: Tomashenko, Natalia, et al.
Pubblicazione: (2026)
di: Tomashenko, Natalia, et al.
Pubblicazione: (2026)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
di: Du, Chenpeng, et al.
Pubblicazione: (2021)
di: Du, Chenpeng, et al.
Pubblicazione: (2021)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
Zero-shot Voice Conversion with Diffusion Transformers
di: Liu, Songting
Pubblicazione: (2024)
di: Liu, Songting
Pubblicazione: (2024)
Self-supervised Reflective Learning through Self-distillation and Online Clustering for Speaker Representation Learning
di: Cai, Danwei, et al.
Pubblicazione: (2024)
di: Cai, Danwei, et al.
Pubblicazione: (2024)
DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
di: Wang, Qing, et al.
Pubblicazione: (2025)
di: Wang, Qing, et al.
Pubblicazione: (2025)
Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR
di: Tan, Kai, et al.
Pubblicazione: (2026)
di: Tan, Kai, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2025) -
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025) -
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
di: Lee, Philip H., et al.
Pubblicazione: (2024) -
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024) -
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
di: Du, Zongyang, et al.
Pubblicazione: (2025)