Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Wirth, Manuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
von: Chia-Pei, et al.
Veröffentlicht: (2026)
von: Chia-Pei, et al.
Veröffentlicht: (2026)
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
von: Debi, Tanusree, et al.
Veröffentlicht: (2026)
von: Debi, Tanusree, et al.
Veröffentlicht: (2026)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
Defending against Indirect Prompt Injection by Instruction Detection
von: Wen, Tongyu, et al.
Veröffentlicht: (2025)
von: Wen, Tongyu, et al.
Veröffentlicht: (2025)
Red Teaming Large Reasoning Models
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
von: Chang, Hongyan, et al.
Veröffentlicht: (2026)
von: Chang, Hongyan, et al.
Veröffentlicht: (2026)
QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
von: Xie, Yuchong, et al.
Veröffentlicht: (2025)
von: Xie, Yuchong, et al.
Veröffentlicht: (2025)
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
von: Zhao, Lei, et al.
Veröffentlicht: (2026)
von: Zhao, Lei, et al.
Veröffentlicht: (2026)
ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
von: Wang, Che, et al.
Veröffentlicht: (2026)
von: Wang, Che, et al.
Veröffentlicht: (2026)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
von: Zhu, Kaijie, et al.
Veröffentlicht: (2025)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2025)
Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking
von: Chen, Junxi, et al.
Veröffentlicht: (2025)
von: Chen, Junxi, et al.
Veröffentlicht: (2025)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
von: Wang, Che, et al.
Veröffentlicht: (2026)
von: Wang, Che, et al.
Veröffentlicht: (2026)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
von: Johnson, Sam, et al.
Veröffentlicht: (2025)
von: Johnson, Sam, et al.
Veröffentlicht: (2025)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
von: Zhao, Wei, et al.
Veröffentlicht: (2026)
von: Zhao, Wei, et al.
Veröffentlicht: (2026)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
von: Dziemian, Mateusz, et al.
Veröffentlicht: (2026)
von: Dziemian, Mateusz, et al.
Veröffentlicht: (2026)
The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense
von: Lan, Qianlong, et al.
Veröffentlicht: (2026)
von: Lan, Qianlong, et al.
Veröffentlicht: (2026)
Red Teaming AI Red Teaming
von: Majumdar, Subhabrata, et al.
Veröffentlicht: (2025)
von: Majumdar, Subhabrata, et al.
Veröffentlicht: (2025)
Bypassing Prompt Injection Detectors through Evasive Injections
von: Rahman, Md Jahedur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Jahedur, et al.
Veröffentlicht: (2026)
RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
von: Liu, Mingrui, et al.
Veröffentlicht: (2026)
von: Liu, Mingrui, et al.
Veröffentlicht: (2026)
PromptLocate: Localizing Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
Defeating Prompt Injections by Design
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2025)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2025)
Evaluation of Prompt Injection Defenses in Large Language Models
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
von: Deep, Priyal, et al.
Veröffentlicht: (2026)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
von: An, Hengyu, et al.
Veröffentlicht: (2025)
von: An, Hengyu, et al.
Veröffentlicht: (2025)
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
von: Sahoo, Devanshu, et al.
Veröffentlicht: (2025)
von: Sahoo, Devanshu, et al.
Veröffentlicht: (2025)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans
von: Faruque, Md Omar, et al.
Veröffentlicht: (2024)
von: Faruque, Md Omar, et al.
Veröffentlicht: (2024)
TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
PromptArmor: Simple yet Effective Prompt Injection Defenses
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
A Red Teaming Roadmap Towards System-Level Safety
von: Wang, Zifan, et al.
Veröffentlicht: (2025)
von: Wang, Zifan, et al.
Veröffentlicht: (2025)
Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing
von: Bhatt, Manish, et al.
Veröffentlicht: (2025)
von: Bhatt, Manish, et al.
Veröffentlicht: (2025)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
von: Cui, Yu, et al.
Veröffentlicht: (2025)
von: Cui, Yu, et al.
Veröffentlicht: (2025)
Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models
von: Ganiuly, Daniyal, et al.
Veröffentlicht: (2025)
von: Ganiuly, Daniyal, et al.
Veröffentlicht: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
von: Zhou, Zhaojiacheng
Veröffentlicht: (2026)
von: Zhou, Zhaojiacheng
Veröffentlicht: (2026)
A Critical Evaluation of Defenses against Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
CourtGuard: A Local, Multiagent Prompt Injection Classifier
von: Wu, Isaac, et al.
Veröffentlicht: (2025)
von: Wu, Isaac, et al.
Veröffentlicht: (2025)
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
von: Ou, Haoran, et al.
Veröffentlicht: (2025)
von: Ou, Haoran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
von: Syros, Georgios, et al.
Veröffentlicht: (2026) -
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
von: Chia-Pei, et al.
Veröffentlicht: (2026) -
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
von: Debi, Tanusree, et al.
Veröffentlicht: (2026) -
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
von: Wang, Zhun, et al.
Veröffentlicht: (2025) -
Defending against Indirect Prompt Injection by Instruction Detection
von: Wen, Tongyu, et al.
Veröffentlicht: (2025)