BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911720519237632 |
|---|---|
| author | Luo, Xuan Wang, Yue Tu, Geng Li, Jing Xu, Ruifeng |
| author_facet | Luo, Xuan Wang, Yue Tu, Geng Li, Jing Xu, Ruifeng |
| contents | In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_27110 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning Luo, Xuan Wang, Yue Tu, Geng Li, Jing Xu, Ruifeng Cryptography and Security Computation and Language In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering. |
| title | BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning |
| topic | Cryptography and Security Computation and Language |
| url | https://arxiv.org/abs/2605.27110 |