Why disentanglement-based speaker anonymization systems fail at preserving emotions?
Fuente:
arXiv
Saved in:
| Main Authors: | Gaznepoglu, Ünal Ege, Peters, Nils |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
by: Ueda, Lucas H., et al.
Published: (2026)
by: Ueda, Lucas H., et al.
Published: (2026)
Hierarchical speaker representation for target speaker extraction
by: He, Shulin, et al.
Published: (2022)
by: He, Shulin, et al.
Published: (2022)
Text adaptation for speaker verification with speaker-text factorized embeddings
by: Yang, Yexin, et al.
Published: (2025)
by: Yang, Yexin, et al.
Published: (2025)
Improving curriculum learning for target speaker extraction with synthetic speakers
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
Improving speaker verification robustness with synthetic emotional utterances
by: Koditala, Nikhil Kumar, et al.
Published: (2024)
by: Koditala, Nikhil Kumar, et al.
Published: (2024)
Visual-based spatial audio generation system for multi-speaker environments
by: Liu, Xiaojing, et al.
Published: (2025)
by: Liu, Xiaojing, et al.
Published: (2025)
Speaker anonymization using neural audio codec language models
by: Panariello, Michele, et al.
Published: (2023)
by: Panariello, Michele, et al.
Published: (2023)
MBCodec:Thorough disentangle for high-fidelity audio compression
by: Zhang, Ruonan, et al.
Published: (2025)
by: Zhang, Ruonan, et al.
Published: (2025)
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
by: Cord-Landwehr, Tobias, et al.
Published: (2025)
by: Cord-Landwehr, Tobias, et al.
Published: (2025)
How phonemes contribute to deep speaker models?
by: Li, Pengqi, et al.
Published: (2024)
by: Li, Pengqi, et al.
Published: (2024)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
by: Ueda, Lucas H., et al.
Published: (2024)
by: Ueda, Lucas H., et al.
Published: (2024)
FreeCodec: A disentangled neural speech codec with fewer tokens
by: Zheng, Youqiang, et al.
Published: (2024)
by: Zheng, Youqiang, et al.
Published: (2024)
VoxATtack: A Multimodal Attack on Voice Anonymization Systems
by: Aloradi, Ahmad, et al.
Published: (2025)
by: Aloradi, Ahmad, et al.
Published: (2025)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
by: Han, Runduo, et al.
Published: (2024)
by: Han, Runduo, et al.
Published: (2024)
The importance of spatial and spectral information in multiple speaker tracking
by: Beit-On, Hanan, et al.
Published: (2024)
by: Beit-On, Hanan, et al.
Published: (2024)
On the influence of language similarity in non-target speaker verification trials
by: Reuter, Paul M., et al.
Published: (2025)
by: Reuter, Paul M., et al.
Published: (2025)
Audio-visual child-adult speaker classification in dyadic interactions
by: Xu, Anfeng, et al.
Published: (2023)
by: Xu, Anfeng, et al.
Published: (2023)
Spoken language change detection inspired by speaker change detection
by: Mishra, Jagabandhu, et al.
Published: (2023)
by: Mishra, Jagabandhu, et al.
Published: (2023)
Contrastive Loss Based Frame-wise Feature disentanglement for Polyphonic Sound Event Detection
by: Guan, Yadong, et al.
Published: (2024)
by: Guan, Yadong, et al.
Published: (2024)
Spectral or spatial? Leveraging both for speaker extraction in challenging data conditions
by: Eisenberg, Aviad, et al.
Published: (2025)
by: Eisenberg, Aviad, et al.
Published: (2025)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
by: Murata, Masato, et al.
Published: (2025)
by: Murata, Masato, et al.
Published: (2025)
Gradient weighting for speaker verification in extremely low Signal-to-Noise Ratio
by: Ma, Yi, et al.
Published: (2024)
by: Ma, Yi, et al.
Published: (2024)
HeightCeleb - an enrichment of VoxCeleb dataset with speaker height information
by: Kacprzak, Stanisław, et al.
Published: (2024)
by: Kacprzak, Stanisław, et al.
Published: (2024)
Improving fairness in speaker verification via Group-adapted Fusion Network
by: Shen, Hua, et al.
Published: (2022)
by: Shen, Hua, et al.
Published: (2022)
A framework of text-dependent speaker verification for chinese numerical string corpus
by: Zheng, Litong, et al.
Published: (2024)
by: Zheng, Litong, et al.
Published: (2024)
EEND-M2F: Masked-attention mask transformers for speaker diarization
by: Härkönen, Marc, et al.
Published: (2024)
by: Härkönen, Marc, et al.
Published: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
by: Liu, Xueyu, et al.
Published: (2024)
by: Liu, Xueyu, et al.
Published: (2024)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
by: Zhang, Yiru, et al.
Published: (2025)
by: Zhang, Yiru, et al.
Published: (2025)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
by: Kunešová, Marie, et al.
Published: (2025)
by: Kunešová, Marie, et al.
Published: (2025)
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
by: Sun, Chang, et al.
Published: (2024)
by: Sun, Chang, et al.
Published: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
by: Grossman, Raymond, et al.
Published: (2025)
by: Grossman, Raymond, et al.
Published: (2025)
You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks
by: Gaznepoglu, Ünal Ege, et al.
Published: (2025)
by: Gaznepoglu, Ünal Ege, et al.
Published: (2025)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
by: Leygue, Tahitoa, et al.
Published: (2025)
by: Leygue, Tahitoa, et al.
Published: (2025)
A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
by: Wang, Xiaopeng, et al.
Published: (2024)
by: Wang, Xiaopeng, et al.
Published: (2024)
Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
by: Zhu, Haolin, et al.
Published: (2024)
by: Zhu, Haolin, et al.
Published: (2024)
Clustering-based hard negative sampling for supervised contrastive speaker verification
by: Masztalski, Piotr, et al.
Published: (2025)
by: Masztalski, Piotr, et al.
Published: (2025)
On the calibration of powerset speaker diarization models
by: Plaquet, Alexis, et al.
Published: (2024)
by: Plaquet, Alexis, et al.
Published: (2024)
A Benchmark for Multi-speaker Anonymization
by: Miao, Xiaoxiao, et al.
Published: (2024)
by: Miao, Xiaoxiao, et al.
Published: (2024)
SCDiar: a streaming diarization system based on speaker change detection and speech recognition
by: Zheng, Naijun, et al.
Published: (2025)
by: Zheng, Naijun, et al.
Published: (2025)
Towards interpretable emotion recognition: Identifying key features with machine learning
by: Kaloga, Yacouba, et al.
Published: (2025)
by: Kaloga, Yacouba, et al.
Published: (2025)
Similar Items
-
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
by: Ueda, Lucas H., et al.
Published: (2026) -
Hierarchical speaker representation for target speaker extraction
by: He, Shulin, et al.
Published: (2022) -
Text adaptation for speaker verification with speaker-text factorized embeddings
by: Yang, Yexin, et al.
Published: (2025) -
Improving curriculum learning for target speaker extraction with synthetic speakers
by: Liu, Yun, et al.
Published: (2024) -
Improving speaker verification robustness with synthetic emotional utterances
by: Koditala, Nikhil Kumar, et al.
Published: (2024)