Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Ze, Shi, Yao, Xu, Yunfei, Li, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
Vclip: Face-based Speaker Generation by Face-voice Association Learning
von: Shi, Yao, et al.
Veröffentlicht: (2026)
von: Shi, Yao, et al.
Veröffentlicht: (2026)
Debatts: Zero-Shot Debating Text-to-Speech Synthesis
von: Huang, Yiqiao, et al.
Veröffentlicht: (2024)
von: Huang, Yiqiao, et al.
Veröffentlicht: (2024)
Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems
von: Joshi, Sonal, et al.
Veröffentlicht: (2021)
von: Joshi, Sonal, et al.
Veröffentlicht: (2021)
Adversarial Attacks and Defenses for Speech Recognition Systems
von: Żelasko, Piotr, et al.
Veröffentlicht: (2021)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2021)
Over-the-Air Adversarial Attack Detection: from Datasets to Defenses
von: Wang, Li, et al.
Veröffentlicht: (2025)
von: Wang, Li, et al.
Veröffentlicht: (2025)
AdvSV: An Over-the-Air Adversarial Attack Dataset for Speaker Verification
von: Wang, Li, et al.
Veröffentlicht: (2023)
von: Wang, Li, et al.
Veröffentlicht: (2023)
DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
von: Yang, Mu, et al.
Veröffentlicht: (2026)
von: Yang, Mu, et al.
Veröffentlicht: (2026)
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
von: Li, Haoyu, et al.
Veröffentlicht: (2025)
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
von: Li, Ze, et al.
Veröffentlicht: (2026)
von: Li, Ze, et al.
Veröffentlicht: (2026)
Zero-Shot Text-to-Speech from Continuous Text Streams
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
von: Li, Ze, et al.
Veröffentlicht: (2025)
von: Li, Ze, et al.
Veröffentlicht: (2025)
Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance
von: Shi, Runwu, et al.
Veröffentlicht: (2025)
von: Shi, Runwu, et al.
Veröffentlicht: (2025)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition Systems
von: Fang, Zheng, et al.
Veröffentlicht: (2024)
von: Fang, Zheng, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
The Database and Benchmark for the Source Speaker Tracing Challenge 2024
von: Li, Ze, et al.
Veröffentlicht: (2024)
von: Li, Ze, et al.
Veröffentlicht: (2024)
Xi+: Uncertainty Supervision for Robust Speaker Embedding
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
von: Feng, Fuyuan, et al.
Veröffentlicht: (2026)
von: Feng, Fuyuan, et al.
Veröffentlicht: (2026)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
von: Li, Junjie, et al.
Veröffentlicht: (2023)
von: Li, Junjie, et al.
Veröffentlicht: (2023)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
Cochleagram-based Noise Adapted Speaker Identification System for Distorted Speech
von: Ahmed, Sabbir, et al.
Veröffentlicht: (2025)
von: Ahmed, Sabbir, et al.
Veröffentlicht: (2025)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
von: Lin, Yuke, et al.
Veröffentlicht: (2025) -
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
von: Zhang, Bowen, et al.
Veröffentlicht: (2025) -
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
von: Lin, Yuke, et al.
Veröffentlicht: (2025) -
Vclip: Face-based Speaker Generation by Face-voice Association Learning
von: Shi, Yao, et al.
Veröffentlicht: (2026) -
Debatts: Zero-Shot Debating Text-to-Speech Synthesis
von: Huang, Yiqiao, et al.
Veröffentlicht: (2024)