Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schmotz, David, Abdelnabi, Sahar, Andriushchenko, Maksym |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
von: Schmotz, David, et al.
Veröffentlicht: (2026)
von: Schmotz, David, et al.
Veröffentlicht: (2026)
Decomposing and Measuring Evaluation Awareness
von: Li, Changling, et al.
Veröffentlicht: (2026)
von: Li, Changling, et al.
Veröffentlicht: (2026)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
AI Agents May Always Fall for Prompt Injections
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals
von: Qin, Jeremy, et al.
Veröffentlicht: (2026)
von: Qin, Jeremy, et al.
Veröffentlicht: (2026)
Does Refusal Training in LLMs Generalize to the Past Tense?
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Context-Aware Reasoning On Parametric Knowledge for Inferring Causal Variables
von: Sheth, Ivaxi, et al.
Veröffentlicht: (2024)
von: Sheth, Ivaxi, et al.
Veröffentlicht: (2024)
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
von: Freeman, Joshua, et al.
Veröffentlicht: (2024)
von: Freeman, Joshua, et al.
Veröffentlicht: (2024)
Watch Out! Simple Horizontal Class Backdoor Can Trivially Evade Defense
von: Ma, Hua, et al.
Veröffentlicht: (2023)
von: Ma, Hua, et al.
Veröffentlicht: (2023)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
von: Salem, Ahmed, et al.
Veröffentlicht: (2026)
von: Salem, Ahmed, et al.
Veröffentlicht: (2026)
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
von: Talokar, Nivya, et al.
Veröffentlicht: (2026)
von: Talokar, Nivya, et al.
Veröffentlicht: (2026)
Why Do We Need Weight Decay in Modern Deep Learning?
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
Critical Influence of Overparameterization on Sharpness-aware Minimization
von: Shin, Sungbin, et al.
Veröffentlicht: (2023)
von: Shin, Sungbin, et al.
Veröffentlicht: (2023)
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
von: Wehner, Jan, et al.
Veröffentlicht: (2025)
von: Wehner, Jan, et al.
Veröffentlicht: (2025)
Layer-wise Linear Mode Connectivity
von: Adilova, Linara, et al.
Veröffentlicht: (2023)
von: Adilova, Linara, et al.
Veröffentlicht: (2023)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
von: Kuntz, Thomas, et al.
Veröffentlicht: (2025)
von: Kuntz, Thomas, et al.
Veröffentlicht: (2025)
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
von: Rose, Aaron, et al.
Veröffentlicht: (2026)
von: Rose, Aaron, et al.
Veröffentlicht: (2026)
No More, No Less: Task Alignment in Terminal Agents
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
von: Rank, Ben, et al.
Veröffentlicht: (2026)
von: Rank, Ben, et al.
Veröffentlicht: (2026)
Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?
von: Zverev, Egor, et al.
Veröffentlicht: (2024)
von: Zverev, Egor, et al.
Veröffentlicht: (2024)
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2023)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2023)
Design Patterns for Securing LLM Agents against Prompt Injections
von: Beurer-Kellner, Luca, et al.
Veröffentlicht: (2025)
von: Beurer-Kellner, Luca, et al.
Veröffentlicht: (2025)
FutureSim: Replaying World Events to Evaluate Adaptive Agents
von: Goel, Shashwat, et al.
Veröffentlicht: (2026)
von: Goel, Shashwat, et al.
Veröffentlicht: (2026)
Evasion of IoT Malware Detection via Dummy Code Injection
von: Zargarzadeh, Sahar, et al.
Veröffentlicht: (2026)
von: Zargarzadeh, Sahar, et al.
Veröffentlicht: (2026)
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
von: Hossain, S M Asif, et al.
Veröffentlicht: (2025)
von: Hossain, S M Asif, et al.
Veröffentlicht: (2025)
Towards Realistic Class-Incremental Learning with Free-Flow Increments
von: Xu, Zhiming, et al.
Veröffentlicht: (2026)
von: Xu, Zhiming, et al.
Veröffentlicht: (2026)
Preventing Prompt Injection with Type-Directed Privilege Separation
von: Jacob, Dennis, et al.
Veröffentlicht: (2025)
von: Jacob, Dennis, et al.
Veröffentlicht: (2025)
IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
von: Terekhov, Mikhail, et al.
Veröffentlicht: (2025)
von: Terekhov, Mikhail, et al.
Veröffentlicht: (2025)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
WebInject: Prompt Injection Attack to Web Agents
von: Wang, Xilong, et al.
Veröffentlicht: (2025)
von: Wang, Xilong, et al.
Veröffentlicht: (2025)
Logic layer Prompt Control Injection (LPCI): A Novel Security Vulnerability Class in Agentic Systems
von: Atta, Hammad, et al.
Veröffentlicht: (2025)
von: Atta, Hammad, et al.
Veröffentlicht: (2025)
SecAlign: Defending Against Prompt Injection with Preference Optimization
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
von: Zhang, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Kaiyuan, et al.
Veröffentlicht: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
von: Schmotz, David, et al.
Veröffentlicht: (2026) -
Decomposing and Measuring Evaluation Awareness
von: Li, Changling, et al.
Veröffentlicht: (2026) -
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024) -
AI Agents May Always Fall for Prompt Injections
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026) -
QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals
von: Qin, Jeremy, et al.
Veröffentlicht: (2026)