Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Shuyi, Suri, Anshuman, Oprea, Alina, Tan, Cheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914504095301632
author Lin, Shuyi
Suri, Anshuman
Oprea, Alina
Tan, Cheng
author_facet Lin, Shuyi
Suri, Anshuman
Oprea, Alina
Tan, Cheng
contents As large language models (LLMs) become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their vulnerability to jailbreak attacks presents a critical security gap. We introduce the jailbreak oracle problem: given a model, prompt, and decoding strategy, determine whether a jailbreak response can be generated with likelihood exceeding a specified threshold. This formalization enables a principled study of jailbreak vulnerabilities. Answering the jailbreak oracle problem poses significant computational challenges, as the search space grows exponentially with response length. We present Boa, the first system designed for efficiently solving the jailbreak oracle problem. Boa employs a two-phase search strategy: (1) breadth-first sampling to identify easily accessible jailbreaks, followed by (2) depth-first priority search guided by fine-grained safety scores to systematically explore promising yet low-probability paths. Boa enables rigorous security assessments including systematic defense evaluation, standardized comparison of red team attacks, and model certification under extreme adversarial conditions. Code is available at https://github.com/shuyilinn/BOA/tree/mlsys2026ae
format Preprint
id arxiv_https___arxiv_org_abs_2506_17299
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
Lin, Shuyi
Suri, Anshuman
Oprea, Alina
Tan, Cheng
Cryptography and Security
Artificial Intelligence
Machine Learning
As large language models (LLMs) become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their vulnerability to jailbreak attacks presents a critical security gap. We introduce the jailbreak oracle problem: given a model, prompt, and decoding strategy, determine whether a jailbreak response can be generated with likelihood exceeding a specified threshold. This formalization enables a principled study of jailbreak vulnerabilities. Answering the jailbreak oracle problem poses significant computational challenges, as the search space grows exponentially with response length. We present Boa, the first system designed for efficiently solving the jailbreak oracle problem. Boa employs a two-phase search strategy: (1) breadth-first sampling to identify easily accessible jailbreaks, followed by (2) depth-first priority search guided by fine-grained safety scores to systematically explore promising yet low-probability paths. Boa enables rigorous security assessments including systematic defense evaluation, standardized comparison of red team attacks, and model certification under extreme adversarial conditions. Code is available at https://github.com/shuyilinn/BOA/tree/mlsys2026ae
title Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.17299