Synthetic Voices, Real Threats: Evaluating Large Text-to-Speech Models in Generating Harmful Audio
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Guangke, Wang, Yuhui, Ji, Shouling, Luo, Xiapu, Wang, Ting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers
by: Chen, Guangke, et al.
Published: (2024)
by: Chen, Guangke, et al.
Published: (2024)
TriniMark: A Robust Generative Speech Watermarking Method for Trinity-Level Traceability
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
VocalCrypt: Novel Active Defense Against Deepfake Voice Based on Masking Effect
by: Fei, Qingyuan, et al.
Published: (2025)
by: Fei, Qingyuan, et al.
Published: (2025)
A Survey on Speech Deepfake Detection
by: Li, Menglu, et al.
Published: (2024)
by: Li, Menglu, et al.
Published: (2024)
IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection
by: Zhu, Jiajie, et al.
Published: (2026)
by: Zhu, Jiajie, et al.
Published: (2026)
EveGuard: Defeating Vibration-based Side-Channel Eavesdropping with Audio Adversarial Perturbations
by: Chang, Jung-Woo, et al.
Published: (2024)
by: Chang, Jung-Woo, et al.
Published: (2024)
Every Breath You Don't Take: Deepfake Speech Detection Using Breath
by: Layton, Seth, et al.
Published: (2024)
by: Layton, Seth, et al.
Published: (2024)
SafeEar: Content Privacy-Preserving Audio Deepfake Detection
by: Li, Xinfeng, et al.
Published: (2024)
by: Li, Xinfeng, et al.
Published: (2024)
Evaluating Synthetic Command Attacks on Smart Voice Assistants
by: He, Zhengxian, et al.
Published: (2024)
by: He, Zhengxian, et al.
Published: (2024)
Hindi audio-video-Deepfake (HAV-DF): A Hindi language-based Audio-video Deepfake Dataset
by: Kaur, Sukhandeep, et al.
Published: (2024)
by: Kaur, Sukhandeep, et al.
Published: (2024)
Cross-Technology Generalization in Synthesized Speech Detection: Evaluating AST Models with Modern Voice Generators
by: Ustinov, Andrew, et al.
Published: (2025)
by: Ustinov, Andrew, et al.
Published: (2025)
A Practical Survey on Emerging Threats from AI-driven Voice Attacks: How Vulnerable are Commercial Voice Control Systems?
by: Wang, Yuanda, et al.
Published: (2023)
by: Wang, Yuanda, et al.
Published: (2023)
Audio-Visual Deepfake Detection With Local Temporal Inconsistencies
by: Astrid, Marcella, et al.
Published: (2025)
by: Astrid, Marcella, et al.
Published: (2025)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
by: Niu, Xinlei, et al.
Published: (2024)
by: Niu, Xinlei, et al.
Published: (2024)
ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
by: Lin, Jinwei
Published: (2024)
by: Lin, Jinwei
Published: (2024)
Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech
by: Huang, Kunyang, et al.
Published: (2025)
by: Huang, Kunyang, et al.
Published: (2025)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
by: Shi, Jiatong, et al.
Published: (2024)
by: Shi, Jiatong, et al.
Published: (2024)
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
by: Mao, Xutao, et al.
Published: (2025)
by: Mao, Xutao, et al.
Published: (2025)
IDEAW: Robust Neural Audio Watermarking with Invertible Dual-Embedding
by: Li, Pengcheng, et al.
Published: (2024)
by: Li, Pengcheng, et al.
Published: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
by: Liu, Qianhui, et al.
Published: (2024)
by: Liu, Qianhui, et al.
Published: (2024)
SuperEar: Eavesdropping on Mobile Voice Calls via Stealthy Acoustic Metamaterials
by: Ning, Zhiyuan, et al.
Published: (2025)
by: Ning, Zhiyuan, et al.
Published: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
by: Yu, Fan, et al.
Published: (2024)
by: Yu, Fan, et al.
Published: (2024)
Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model
by: Chen, Gehui, et al.
Published: (2024)
by: Chen, Gehui, et al.
Published: (2024)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
by: Liu, Haohe, et al.
Published: (2024)
by: Liu, Haohe, et al.
Published: (2024)
The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents
by: Wang, Lixu, et al.
Published: (2025)
by: Wang, Lixu, et al.
Published: (2025)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
by: Li, Fengjin, et al.
Published: (2025)
by: Li, Fengjin, et al.
Published: (2025)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
by: Tian, Wenjie, et al.
Published: (2025)
by: Tian, Wenjie, et al.
Published: (2025)
LoVA: Long-form Video-to-Audio Generation
by: Cheng, Xin, et al.
Published: (2024)
by: Cheng, Xin, et al.
Published: (2024)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
by: Pan, Tianrui, et al.
Published: (2024)
by: Pan, Tianrui, et al.
Published: (2024)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
by: Huang, Zhiqi, et al.
Published: (2024)
by: Huang, Zhiqi, et al.
Published: (2024)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
by: Guo, Zixin, et al.
Published: (2024)
by: Guo, Zixin, et al.
Published: (2024)
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
by: Chen, Guangke, et al.
Published: (2025)
by: Chen, Guangke, et al.
Published: (2025)
Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer
by: Jin, Weifei, et al.
Published: (2024)
by: Jin, Weifei, et al.
Published: (2024)
Audio-Visual Speech Separation via Bottleneck Iterative Network
by: Zhang, Sidong, et al.
Published: (2025)
by: Zhang, Sidong, et al.
Published: (2025)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
by: Lin, Yueqian, et al.
Published: (2025)
by: Lin, Yueqian, et al.
Published: (2025)
A Preliminary Case Study on Long-Form In-the-Wild Audio Spoofing Detection
by: Liu, Xuechen, et al.
Published: (2024)
by: Liu, Xuechen, et al.
Published: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
by: Yuan, Yi, et al.
Published: (2024)
by: Yuan, Yi, et al.
Published: (2024)
V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation
by: Chan, Nolan, et al.
Published: (2026)
by: Chan, Nolan, et al.
Published: (2026)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
by: Biyani, Ishan D., et al.
Published: (2025)
by: Biyani, Ishan D., et al.
Published: (2025)
Similar Items
-
SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers
by: Chen, Guangke, et al.
Published: (2024) -
TriniMark: A Robust Generative Speech Watermarking Method for Trinity-Level Traceability
by: Li, Yue, et al.
Published: (2025) -
VocalCrypt: Novel Active Defense Against Deepfake Voice Based on Masking Effect
by: Fei, Qingyuan, et al.
Published: (2025) -
A Survey on Speech Deepfake Detection
by: Li, Menglu, et al.
Published: (2024) -
IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection
by: Zhu, Jiajie, et al.
Published: (2026)