Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Sizhe, Zharmagambetov, Arman, Wagner, David, Guo, Chuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SecAlign: Defending Against Prompt Injection with Preference Optimization
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
von: Evtimov, Ivan, et al.
Veröffentlicht: (2025)
von: Evtimov, Ivan, et al.
Veröffentlicht: (2025)
Securing AI Agents Against Prompt Injection Attacks
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
von: Wang, Zhilong, et al.
Veröffentlicht: (2025)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
von: Suo, Xuchen
Veröffentlicht: (2024)
von: Suo, Xuchen
Veröffentlicht: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
von: Zhu, Kaijie, et al.
Veröffentlicht: (2025)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2025)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
von: Zhao, Wei, et al.
Veröffentlicht: (2026)
von: Zhao, Wei, et al.
Veröffentlicht: (2026)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
von: Zhong, Peter Yong, et al.
Veröffentlicht: (2025)
von: Zhong, Peter Yong, et al.
Veröffentlicht: (2025)
StruQ: Defending Against Prompt Injection with Structured Queries
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
Defending Against Prompt Injection with DataFilter
von: Wang, Yizhu, et al.
Veröffentlicht: (2025)
von: Wang, Yizhu, et al.
Veröffentlicht: (2025)
SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
von: Zhang, Jiawen, et al.
Veröffentlicht: (2025)
von: Zhang, Jiawen, et al.
Veröffentlicht: (2025)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
von: Piet, Julien, et al.
Veröffentlicht: (2023)
von: Piet, Julien, et al.
Veröffentlicht: (2023)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
PromptLocate: Localizing Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
Defending Against Prompt Injection With a Few DefensiveTokens
von: Chen, Sizhe, et al.
Veröffentlicht: (2025)
von: Chen, Sizhe, et al.
Veröffentlicht: (2025)
Encrypted Prompt: Securing LLM Applications Against Unauthorized Actions
von: Chan, Shih-Han
Veröffentlicht: (2025)
von: Chan, Shih-Han
Veröffentlicht: (2025)
Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
von: Pasquini, Dario, et al.
Veröffentlicht: (2024)
von: Pasquini, Dario, et al.
Veröffentlicht: (2024)
PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems
von: Wang, Haozhen, et al.
Veröffentlicht: (2026)
von: Wang, Haozhen, et al.
Veröffentlicht: (2026)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
Design and Implementation of a Secure RAG-Enhanced AI Chatbot for Smart Tourism Customer Service: Defending Against Prompt Injection Attacks -- A Case Study of Hsinchu, Taiwan
von: Shih, Yu-Kai, et al.
Veröffentlicht: (2025)
von: Shih, Yu-Kai, et al.
Veröffentlicht: (2025)
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
von: Li, Zongze, et al.
Veröffentlicht: (2025)
von: Li, Zongze, et al.
Veröffentlicht: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
von: Johnson, Sam, et al.
Veröffentlicht: (2025)
von: Johnson, Sam, et al.
Veröffentlicht: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
von: Cao, Tri, et al.
Veröffentlicht: (2026)
von: Cao, Tri, et al.
Veröffentlicht: (2026)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
von: Panebianco, Francesco, et al.
Veröffentlicht: (2025)
von: Panebianco, Francesco, et al.
Veröffentlicht: (2025)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
von: Cui, Yu, et al.
Veröffentlicht: (2025)
von: Cui, Yu, et al.
Veröffentlicht: (2025)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
von: Yeo, Andrew, et al.
Veröffentlicht: (2025)
von: Yeo, Andrew, et al.
Veröffentlicht: (2025)
The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense
von: Lan, Qianlong, et al.
Veröffentlicht: (2026)
von: Lan, Qianlong, et al.
Veröffentlicht: (2026)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
von: Chen, Qirui, et al.
Veröffentlicht: (2026)
von: Chen, Qirui, et al.
Veröffentlicht: (2026)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
Defenses Against Prompt Attacks Learn Surface Heuristics
von: Li, Shawn, et al.
Veröffentlicht: (2026)
von: Li, Shawn, et al.
Veröffentlicht: (2026)
Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SecAlign: Defending Against Prompt Injection with Preference Optimization
von: Chen, Sizhe, et al.
Veröffentlicht: (2024) -
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
von: Evtimov, Ivan, et al.
Veröffentlicht: (2025) -
Securing AI Agents Against Prompt Injection Attacks
von: Ramakrishnan, Badrinath, et al.
Veröffentlicht: (2025) -
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
von: Wang, Zhilong, et al.
Veröffentlicht: (2025) -
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)