Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xin, Yuan, Chen, Dingfan, Yang, Linyi, Backes, Michael, Zhang, Xiao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909978614300672
author Xin, Yuan
Chen, Dingfan
Yang, Linyi
Backes, Michael
Zhang, Xiao
author_facet Xin, Yuan
Chen, Dingfan
Yang, Linyi
Backes, Michael
Zhang, Xiao
contents As large language models (LLMs) are increasingly deployed, ensuring their safe use is paramount. Jailbreaking, adversarial prompts that bypass model alignment to trigger harmful outputs, present significant risks, with existing studies reporting high success rates in evading common LLMs. However, previous evaluations have focused solely on the models, neglecting the full deployment pipeline, which typically incorporates additional safety mechanisms like content moderation filters. To address this gap, we present the first systematic evaluation of jailbreak attacks targeting LLM safety alignment, assessing their success across the full inference pipeline, including both input and output filtering stages. Our findings yield two key insights: first, nearly all evaluated jailbreak techniques can be detected by at least one safety filter, suggesting that prior assessments may have overestimated the practical success of these attacks; second, while safety filters are effective in detection, there remains room to better balance recall and precision to further optimize protection and user experience. We highlight critical gaps and call for further refinement of detection accuracy and usability in LLM safety systems.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24044
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
Xin, Yuan
Chen, Dingfan
Yang, Linyi
Backes, Michael
Zhang, Xiao
Cryptography and Security
Artificial Intelligence
Computation and Language
I.2.7
As large language models (LLMs) are increasingly deployed, ensuring their safe use is paramount. Jailbreaking, adversarial prompts that bypass model alignment to trigger harmful outputs, present significant risks, with existing studies reporting high success rates in evading common LLMs. However, previous evaluations have focused solely on the models, neglecting the full deployment pipeline, which typically incorporates additional safety mechanisms like content moderation filters. To address this gap, we present the first systematic evaluation of jailbreak attacks targeting LLM safety alignment, assessing their success across the full inference pipeline, including both input and output filtering stages. Our findings yield two key insights: first, nearly all evaluated jailbreak techniques can be detected by at least one safety filter, suggesting that prior assessments may have overestimated the practical success of these attacks; second, while safety filters are effective in detection, there remains room to better balance recall and precision to further optimize protection and user experience. We highlight critical gaps and call for further refinement of detection accuracy and usability in LLM safety systems.
title Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
topic Cryptography and Security
Artificial Intelligence
Computation and Language
I.2.7
url https://arxiv.org/abs/2512.24044