Jailbreaking in the Haystack

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shah, Rishi Rajesh, Wu, Chen Henry, Saxena, Shashwat, Zhong, Ziqian, Robey, Alexander, Raghunathan, Aditi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915603896336384
author Shah, Rishi Rajesh
Wu, Chen Henry
Saxena, Shashwat
Zhong, Ziqian
Robey, Alexander
Raghunathan, Aditi
author_facet Shah, Rishi Rajesh
Wu, Chen Henry
Saxena, Shashwat
Zhong, Ziqian
Robey, Alexander
Raghunathan, Aditi
contents Recent advances in long-context language models (LMs) have enabled million-token inputs, expanding their capabilities across complex tasks like computer-use agents. Yet, the safety implications of these extended contexts remain unclear. To bridge this gap, we introduce NINJA (short for Needle-in-haystack jailbreak attack), a method that jailbreaks aligned LMs by appending benign, model-generated content to harmful user goals. Critical to our method is the observation that the position of harmful goals play an important role in safety. Experiments on standard safety benchmark, HarmBench, show that NINJA significantly increases attack success rates across state-of-the-art open and proprietary models, including LLaMA, Qwen, Mistral, and Gemini. Unlike prior jailbreaking methods, our approach is low-resource, transferable, and less detectable. Moreover, we show that NINJA is compute-optimal -- under a fixed compute budget, increasing context length can outperform increasing the number of trials in best-of-N jailbreak. These findings reveal that even benign long contexts -- when crafted with careful goal positioning -- introduce fundamental vulnerabilities in modern LMs.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04707
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jailbreaking in the Haystack
Shah, Rishi Rajesh
Wu, Chen Henry
Saxena, Shashwat
Zhong, Ziqian
Robey, Alexander
Raghunathan, Aditi
Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
Recent advances in long-context language models (LMs) have enabled million-token inputs, expanding their capabilities across complex tasks like computer-use agents. Yet, the safety implications of these extended contexts remain unclear. To bridge this gap, we introduce NINJA (short for Needle-in-haystack jailbreak attack), a method that jailbreaks aligned LMs by appending benign, model-generated content to harmful user goals. Critical to our method is the observation that the position of harmful goals play an important role in safety. Experiments on standard safety benchmark, HarmBench, show that NINJA significantly increases attack success rates across state-of-the-art open and proprietary models, including LLaMA, Qwen, Mistral, and Gemini. Unlike prior jailbreaking methods, our approach is low-resource, transferable, and less detectable. Moreover, we show that NINJA is compute-optimal -- under a fixed compute budget, increasing context length can outperform increasing the number of trials in best-of-N jailbreak. These findings reveal that even benign long contexts -- when crafted with careful goal positioning -- introduce fundamental vulnerabilities in modern LMs.
title Jailbreaking in the Haystack
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2511.04707