Prompt Injection Attacks in Defended Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Khomsky, Daniil, Maloyan, Narek, Nutfullin, Bulat |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Trojan Detection in Large Language Models: Insights from The Trojan Detection Challenge
par: Maloyan, Narek, et autres
Publié: (2024)
par: Maloyan, Narek, et autres
Publié: (2024)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
par: Maloyan, Narek, et autres
Publié: (2025)
par: Maloyan, Narek, et autres
Publié: (2025)
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
par: Maloyan, Narek, et autres
Publié: (2025)
par: Maloyan, Narek, et autres
Publié: (2025)
Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems
par: Maloyan, Narek, et autres
Publié: (2026)
par: Maloyan, Narek, et autres
Publié: (2026)
Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents
par: Maloyan, Narek, et autres
Publié: (2026)
par: Maloyan, Narek, et autres
Publié: (2026)
Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents
par: Maloyan, Narek, et autres
Publié: (2026)
par: Maloyan, Narek, et autres
Publié: (2026)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
par: Hines, Keegan, et autres
Publié: (2024)
par: Hines, Keegan, et autres
Publié: (2024)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
par: Yi, Jingwei, et autres
Publié: (2023)
par: Yi, Jingwei, et autres
Publié: (2023)
SPML: A DSL for Defending Language Models Against Prompt Attacks
par: Sharma, Reshabh K, et autres
Publié: (2024)
par: Sharma, Reshabh K, et autres
Publié: (2024)
Cognitive Overload Attack:Prompt Injection for Long Context
par: Upadhayay, Bibek, et autres
Publié: (2024)
par: Upadhayay, Bibek, et autres
Publié: (2024)
Defense against Prompt Injection Attacks via Mixture of Encodings
par: Zhang, Ruiyi, et autres
Publié: (2025)
par: Zhang, Ruiyi, et autres
Publié: (2025)
Uncertainty-Aware Evaluation for Vision-Language Models
par: Kostumov, Vasily, et autres
Publié: (2024)
par: Kostumov, Vasily, et autres
Publié: (2024)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
par: Zhou, Andy, et autres
Publié: (2024)
par: Zhou, Andy, et autres
Publié: (2024)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
par: Wang, Jiawen, et autres
Publié: (2025)
par: Wang, Jiawen, et autres
Publié: (2025)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
par: Larionov, Daniil, et autres
Publié: (2025)
par: Larionov, Daniil, et autres
Publié: (2025)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
par: Larionov, Daniil, et autres
Publié: (2024)
par: Larionov, Daniil, et autres
Publié: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
par: Liu, Yupei, et autres
Publié: (2023)
par: Liu, Yupei, et autres
Publié: (2023)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
par: Jia, Feiran, et autres
Publié: (2024)
par: Jia, Feiran, et autres
Publié: (2024)
Scaling Behavior of Machine Translation with Large Language Models under Prompt Injection Attacks
par: Sun, Zhifan, et autres
Publié: (2024)
par: Sun, Zhifan, et autres
Publié: (2024)
Multilingual Hidden Prompt Injection Attacks on LLM-Based Academic Reviewing
par: Theocharopoulos, Panagiotis, et autres
Publié: (2025)
par: Theocharopoulos, Panagiotis, et autres
Publié: (2025)
Defending Against Social Engineering Attacks in the Age of LLMs
par: Ai, Lin, et autres
Publié: (2024)
par: Ai, Lin, et autres
Publié: (2024)
WebInject: Prompt Injection Attack to Web Agents
par: Wang, Xilong, et autres
Publié: (2025)
par: Wang, Xilong, et autres
Publié: (2025)
An Early Categorization of Prompt Injection Attacks on Large Language Models
par: Rossi, Sippo, et autres
Publié: (2024)
par: Rossi, Sippo, et autres
Publié: (2024)
MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs
par: Lee, Junhyeok, et autres
Publié: (2026)
par: Lee, Junhyeok, et autres
Publié: (2026)
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
par: Wang, Peiran, et autres
Publié: (2025)
par: Wang, Peiran, et autres
Publié: (2025)
Goal-guided Generative Prompt Injection Attack on Large Language Models
par: Zhang, Chong, et autres
Publié: (2024)
par: Zhang, Chong, et autres
Publié: (2024)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
par: Wang, Xilong, et autres
Publié: (2026)
par: Wang, Xilong, et autres
Publié: (2026)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
par: Suri, Omar Farooq Khan, et autres
Publié: (2025)
par: Suri, Omar Farooq Khan, et autres
Publié: (2025)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
par: Xin, Yuan, et autres
Publié: (2026)
par: Xin, Yuan, et autres
Publié: (2026)
Fine-tuned Large Language Models (LLMs): Improved Prompt Injection Attacks Detection
par: Rahman, Md Abdur, et autres
Publié: (2024)
par: Rahman, Md Abdur, et autres
Publié: (2024)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
par: Wang, Jiongxiao, et autres
Publié: (2024)
par: Wang, Jiongxiao, et autres
Publié: (2024)
Defending LLMs against Jailbreaking Attacks via Backtranslation
par: Wang, Yihan, et autres
Publié: (2024)
par: Wang, Yihan, et autres
Publié: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
par: Shao, Zedian, et autres
Publié: (2024)
par: Shao, Zedian, et autres
Publié: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
par: Li, Hao, et autres
Publié: (2026)
par: Li, Hao, et autres
Publié: (2026)
Applying Pre-trained Multilingual BERT in Embeddings for Improved Malicious Prompt Injection Attacks Detection
par: Rahman, Md Abdur, et autres
Publié: (2024)
par: Rahman, Md Abdur, et autres
Publié: (2024)
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
par: Ji, Jiabao, et autres
Publié: (2024)
par: Ji, Jiabao, et autres
Publié: (2024)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
par: Zhang, Zhexin, et autres
Publié: (2023)
par: Zhang, Zhexin, et autres
Publié: (2023)
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
par: Li, Yu, et autres
Publié: (2025)
par: Li, Yu, et autres
Publié: (2025)
Defending Against Disinformation Attacks in Open-Domain Question Answering
par: Weller, Orion, et autres
Publié: (2022)
par: Weller, Orion, et autres
Publié: (2022)
Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting
par: Kang, Jingyi, et autres
Publié: (2026)
par: Kang, Jingyi, et autres
Publié: (2026)
Documents similaires
-
Trojan Detection in Large Language Models: Insights from The Trojan Detection Challenge
par: Maloyan, Narek, et autres
Publié: (2024) -
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
par: Maloyan, Narek, et autres
Publié: (2025) -
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
par: Maloyan, Narek, et autres
Publié: (2025) -
Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems
par: Maloyan, Narek, et autres
Publié: (2026) -
Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents
par: Maloyan, Narek, et autres
Publié: (2026)