Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Yanxu, Liu, Peipei, Cui, Tiehan, Liu, Congying, Xing, Mingzhe, You, Datao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914465867366400
author Mao, Yanxu
Liu, Peipei
Cui, Tiehan
Liu, Congying
Xing, Mingzhe
You, Datao
author_facet Mao, Yanxu
Liu, Peipei
Cui, Tiehan
Liu, Congying
Xing, Mingzhe
You, Datao
contents With the widespread application of LLM-based agents across various domains, their complexity has introduced new security threats. Existing red-team methods mostly rely on modifying user prompts, which lack adaptability to new data and may impact the agent's performance. To address the challenge, this paper proposes the JailAgent framework, which completely avoids modifying the user prompt. Specifically, it implicitly manipulates the agent's reasoning trajectory and memory retrieval with three key stages: Trigger Extraction, Reasoning Hijacking, and Constraint Tightening. Through precise trigger identification, real-time adaptive mechanisms, and an optimized objective function, JailAgent demonstrates outstanding performance in cross-model and cross-scenario environments.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05549
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents
Mao, Yanxu
Liu, Peipei
Cui, Tiehan
Liu, Congying
Xing, Mingzhe
You, Datao
Computation and Language
With the widespread application of LLM-based agents across various domains, their complexity has introduced new security threats. Existing red-team methods mostly rely on modifying user prompts, which lack adaptability to new data and may impact the agent's performance. To address the challenge, this paper proposes the JailAgent framework, which completely avoids modifying the user prompt. Specifically, it implicitly manipulates the agent's reasoning trajectory and memory retrieval with three key stages: Trigger Extraction, Reasoning Hijacking, and Constraint Tightening. Through precise trigger identification, real-time adaptive mechanisms, and an optimized objective function, JailAgent demonstrates outstanding performance in cross-model and cross-scenario environments.
title Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents
topic Computation and Language
url https://arxiv.org/abs/2604.05549