SPIRIT: Patching Speech Language Models against Jailbreak Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Djanibekov, Amirbek, Mukhituly, Nurdaulet, Inui, Kentaro, Aldarmaki, Hanan, Lukas, Nils |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dialectal Coverage And Generalization in Arabic Speech Recognition
by: Djanibekov, Amirbek, et al.
Published: (2024)
by: Djanibekov, Amirbek, et al.
Published: (2024)
RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
by: Aldarmaki, Ibrahim, et al.
Published: (2024)
by: Aldarmaki, Ibrahim, et al.
Published: (2024)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
by: Toyin, Hawau Olamide, et al.
Published: (2024)
by: Toyin, Hawau Olamide, et al.
Published: (2024)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
by: Djanibekov, Amirbek, et al.
Published: (2026)
by: Djanibekov, Amirbek, et al.
Published: (2026)
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
by: Chen, Guangke, et al.
Published: (2025)
by: Chen, Guangke, et al.
Published: (2025)
A Systematic Evaluation of Adversarial Attacks against Speech Emotion Recognition Models
by: Facchinetti, Nicolas, et al.
Published: (2024)
by: Facchinetti, Nicolas, et al.
Published: (2024)
"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models
by: Gupta, Isha, et al.
Published: (2025)
by: Gupta, Isha, et al.
Published: (2025)
PatchDSU: Uncertainty Modeling for Out of Distribution Generalization in Keyword Spotting
by: Chernyak, Bronya Roni, et al.
Published: (2025)
by: Chernyak, Bronya Roni, et al.
Published: (2025)
Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?
by: Makarov, Rostislav, et al.
Published: (2025)
by: Makarov, Rostislav, et al.
Published: (2025)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
by: Ji, Shengpeng, et al.
Published: (2023)
by: Ji, Shengpeng, et al.
Published: (2023)
Prompt Amplification and Zero-Shot Late Fusion in Audio-Language Models for Speech Emotion Recognition
by: Kataria, Saurabh, et al.
Published: (2026)
by: Kataria, Saurabh, et al.
Published: (2026)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
by: Wang, Yuancheng, et al.
Published: (2025)
by: Wang, Yuancheng, et al.
Published: (2025)
PALM: Few-Shot Prompt Learning for Audio Language Models
by: Hanif, Asif, et al.
Published: (2024)
by: Hanif, Asif, et al.
Published: (2024)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
by: Hono, Yukiya, et al.
Published: (2023)
by: Hono, Yukiya, et al.
Published: (2023)
Does Language Matter for Early Detection of Parkinson's Disease from Speech?
by: Plantinga, Peter, et al.
Published: (2025)
by: Plantinga, Peter, et al.
Published: (2025)
Discrete-Time Diffusion-Like Models for Speech Synthesis
by: Tan, Xiaozhou, et al.
Published: (2025)
by: Tan, Xiaozhou, et al.
Published: (2025)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
by: Agro, Maha Tufail, et al.
Published: (2025)
by: Agro, Maha Tufail, et al.
Published: (2025)
Clean Label Attacks against SLU Systems
by: Xinyuan, Henry Li, et al.
Published: (2024)
by: Xinyuan, Henry Li, et al.
Published: (2024)
Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets
by: Pešán, Jan, et al.
Published: (2024)
by: Pešán, Jan, et al.
Published: (2024)
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
by: Lovelace, Justin, et al.
Published: (2025)
by: Lovelace, Justin, et al.
Published: (2025)
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
by: Alexey, Protopopov
Published: (2026)
by: Alexey, Protopopov
Published: (2026)
MORE: Multi-Objective Adversarial Attacks on Speech Recognition
by: Gao, Xiaoxue, et al.
Published: (2026)
by: Gao, Xiaoxue, et al.
Published: (2026)
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
by: Narain, Jaya, et al.
Published: (2025)
by: Narain, Jaya, et al.
Published: (2025)
Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
by: Gerazov, Branislav, et al.
Published: (2025)
by: Gerazov, Branislav, et al.
Published: (2025)
From Black Box to Biomarker: Sparse Autoencoders for Interpreting Speech Models of Parkinson's Disease
by: Plantinga, Peter, et al.
Published: (2025)
by: Plantinga, Peter, et al.
Published: (2025)
Investigating the Effects of Diffusion-based Conditional Generative Speech Models Used for Speech Enhancement on Dysarthric Speech
by: Reszka, Joanna, et al.
Published: (2024)
by: Reszka, Joanna, et al.
Published: (2024)
Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity
by: Ulgen, Ismail Rasim, et al.
Published: (2024)
by: Ulgen, Ismail Rasim, et al.
Published: (2024)
Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech
by: de Oliveira, Danilo, et al.
Published: (2024)
by: de Oliveira, Danilo, et al.
Published: (2024)
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
by: Sarkar, Eklavya, et al.
Published: (2025)
by: Sarkar, Eklavya, et al.
Published: (2025)
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
by: Cheng, Hao, et al.
Published: (2025)
by: Cheng, Hao, et al.
Published: (2025)
AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling
by: Kalkhorani, Vahid Ahmadi, et al.
Published: (2024)
by: Kalkhorani, Vahid Ahmadi, et al.
Published: (2024)
Principled Coarse-Grained Acceptance for Speculative Decoding in Speech
by: Yanuka, Moran, et al.
Published: (2025)
by: Yanuka, Moran, et al.
Published: (2025)
Speech Understanding on Tiny Devices with A Learning Cache
by: Benazir, Afsara, et al.
Published: (2023)
by: Benazir, Afsara, et al.
Published: (2023)
Predictive-Generative Drift Decomposition for Speech Enhancement and Separation
by: Richter, Julius, et al.
Published: (2026)
by: Richter, Julius, et al.
Published: (2026)
Robustness of Speech Separation Models for Similar-pitch Speakers
by: Lay, Bunlong, et al.
Published: (2024)
by: Lay, Bunlong, et al.
Published: (2024)
Modeling Overlapped Speech with Shuffles
by: Wiesner, Matthew, et al.
Published: (2026)
by: Wiesner, Matthew, et al.
Published: (2026)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
by: Yamamoto, Ryuichi, et al.
Published: (2024)
by: Yamamoto, Ryuichi, et al.
Published: (2024)
Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
by: Vecino, Biel Tura, et al.
Published: (2025)
by: Vecino, Biel Tura, et al.
Published: (2025)
MaskCycleGAN-based Whisper to Normal Speech Conversion
by: Gupta, K. Rohith, et al.
Published: (2024)
by: Gupta, K. Rohith, et al.
Published: (2024)
Speech to Speech Synthesis for Voice Impersonation
by: Johnson, Bjorn, et al.
Published: (2026)
by: Johnson, Bjorn, et al.
Published: (2026)
Similar Items
-
Dialectal Coverage And Generalization in Arabic Speech Recognition
by: Djanibekov, Amirbek, et al.
Published: (2024) -
RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
by: Aldarmaki, Ibrahim, et al.
Published: (2024) -
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
by: Toyin, Hawau Olamide, et al.
Published: (2024) -
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
by: Djanibekov, Amirbek, et al.
Published: (2026) -
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
by: Chen, Guangke, et al.
Published: (2025)