Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Roh, Jaechul, Houmansadr, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multilingual and Multi-Accent Jailbreaking of Audio LLMs
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
OSLO: One-Shot Label-Only Membership Inference Attacks
by: Peng, Yuefeng, et al.
Published: (2024)
by: Peng, Yuefeng, et al.
Published: (2024)
Codec-Robust Attacks on Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026)
by: Roh, Jaechul, et al.
Published: (2026)
Backdooring Bias ($B^2$) into Stable Diffusion Models
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection
by: Hu, Bin, et al.
Published: (2025)
by: Hu, Bin, et al.
Published: (2025)
OverThink: Slowdown Attacks on Reasoning LLMs
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
Throttling Web Agents Using Reasoning Gates
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech
by: Huang, Kunyang, et al.
Published: (2025)
by: Huang, Kunyang, et al.
Published: (2025)
Breaking Audio Large Language Models by Attacking Only the Encoder: A Universal Targeted Latent-Space Audio Attack
by: Ziv, Roee, et al.
Published: (2025)
by: Ziv, Roee, et al.
Published: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Audio Pirates: Black-box Audio Watermark Removal via Diffusion Priors
by: Yao, Lingfeng, et al.
Published: (2026)
by: Yao, Lingfeng, et al.
Published: (2026)
ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
by: Jin, Weifei, et al.
Published: (2025)
by: Jin, Weifei, et al.
Published: (2025)
Vulnerabilities of Audio-Based Biometric Authentication Systems Against Deepfake Speech Synthesis
by: Hong, Mengze, et al.
Published: (2026)
by: Hong, Mengze, et al.
Published: (2026)
SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
by: Lin, Weilin, et al.
Published: (2025)
by: Lin, Weilin, et al.
Published: (2025)
When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
by: Dingeto, Hiskias, et al.
Published: (2025)
by: Dingeto, Hiskias, et al.
Published: (2025)
MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech
by: Jin, Yutong, et al.
Published: (2026)
by: Jin, Yutong, et al.
Published: (2026)
Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models
by: Wang, Yanyun, et al.
Published: (2026)
by: Wang, Yanyun, et al.
Published: (2026)
SyncGuard: Robust Audio Watermarking Capable of Countering Desynchronization Attacks
by: Gan, Zhenliang, et al.
Published: (2025)
by: Gan, Zhenliang, et al.
Published: (2025)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
PRoADS: Provably Secure and Robust Audio Diffusion Steganography with latent optimization and backward Euler Inversion
by: Yan, YongPeng, et al.
Published: (2026)
by: Yan, YongPeng, et al.
Published: (2026)
Measuring the Robustness of Audio Deepfake Detectors
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
STEP: Detecting Audio Backdoor Attacks via Stability-based Trigger Exposure Profiling
by: Wang, Kun, et al.
Published: (2026)
by: Wang, Kun, et al.
Published: (2026)
TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning
by: Xie, Zhixin, et al.
Published: (2026)
by: Xie, Zhixin, et al.
Published: (2026)
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
by: Chen, Meng, et al.
Published: (2026)
by: Chen, Meng, et al.
Published: (2026)
SilentCipher: Deep Audio Watermarking
by: Singh, Mayank Kumar, et al.
Published: (2024)
by: Singh, Mayank Kumar, et al.
Published: (2024)
Yours or Mine? Overwriting Attacks Against Neural Audio Watermarking
by: Yao, Lingfeng, et al.
Published: (2025)
by: Yao, Lingfeng, et al.
Published: (2025)
Interpretable Temporal Class Activation Representation for Audio Spoofing Detection
by: Li, Menglu, et al.
Published: (2024)
by: Li, Menglu, et al.
Published: (2024)
Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs
by: Ling, Zijian, et al.
Published: (2026)
by: Ling, Zijian, et al.
Published: (2026)
Pitch Imperfect: Detecting Audio Deepfakes Through Acoustic Prosodic Analysis
by: Warren, Kevin, et al.
Published: (2025)
by: Warren, Kevin, et al.
Published: (2025)
One-Class Learning with Adaptive Centroid Shift for Audio Deepfake Detection
by: Kim, Hyun Myung, et al.
Published: (2024)
by: Kim, Hyun Myung, et al.
Published: (2024)
DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition
by: Maurya, Bikrant Bikram Pratap, et al.
Published: (2026)
by: Maurya, Bikrant Bikram Pratap, et al.
Published: (2026)
DASM: Domain-Aware Sharpness Minimization for Multi-Domain Voice Stream Steganalysis
by: Zhou, Pengcheng, et al.
Published: (2026)
by: Zhou, Pengcheng, et al.
Published: (2026)
Invisible Ears at Your Fingertips: Acoustic Eavesdropping via Mouse Sensors
by: Fakih, Mohamad, et al.
Published: (2025)
by: Fakih, Mohamad, et al.
Published: (2025)
Mirage Fools the Ear, Mute Hides the Truth: Precise Targeted Adversarial Attacks on Polyphonic Sound Event Detection Systems
by: Su, Junjie, et al.
Published: (2025)
by: Su, Junjie, et al.
Published: (2025)
Selective Masking Adversarial Attack on Automatic Speech Recognition Systems
by: Fang, Zheng, et al.
Published: (2025)
by: Fang, Zheng, et al.
Published: (2025)
ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
by: Wang, Yuanda, et al.
Published: (2025)
by: Wang, Yuanda, et al.
Published: (2025)
HVAC-EAR: Eavesdropping Human Speech Using HVAC Systems
by: Tamiti, Tarikul Islam, et al.
Published: (2025)
by: Tamiti, Tarikul Islam, et al.
Published: (2025)
A Preliminary Case Study on Long-Form In-the-Wild Audio Spoofing Detection
by: Liu, Xuechen, et al.
Published: (2024)
by: Liu, Xuechen, et al.
Published: (2024)
PITCH: AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response
by: Mittal, Govind, et al.
Published: (2024)
by: Mittal, Govind, et al.
Published: (2024)
AudioMarkBench: Benchmarking Robustness of Audio Watermarking
by: Liu, Hongbin, et al.
Published: (2024)
by: Liu, Hongbin, et al.
Published: (2024)
Similar Items
-
Multilingual and Multi-Accent Jailbreaking of Audio LLMs
by: Roh, Jaechul, et al.
Published: (2025) -
OSLO: One-Shot Label-Only Membership Inference Attacks
by: Peng, Yuefeng, et al.
Published: (2024) -
Codec-Robust Attacks on Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026) -
Backdooring Bias ($B^2$) into Stable Diffusion Models
by: Naseh, Ali, et al.
Published: (2024) -
When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection
by: Hu, Bin, et al.
Published: (2025)