BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Xuan, Wang, Yue, Tu, Geng, Li, Jing, Xu, Ruifeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911720519237632
author Luo, Xuan
Wang, Yue
Tu, Geng
Li, Jing
Xu, Ruifeng
author_facet Luo, Xuan
Wang, Yue
Tu, Geng
Li, Jing
Xu, Ruifeng
contents In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27110
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
Luo, Xuan
Wang, Yue
Tu, Geng
Li, Jing
Xu, Ruifeng
Cryptography and Security
Computation and Language
In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering.
title BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2605.27110