AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Ting-Chun, Hsu, Ching-Yu, Lee, Kuan-Yi, Fu, Chi-An, Lee, Hung-yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916997693964288
author Liu, Ting-Chun
Hsu, Ching-Yu
Lee, Kuan-Yi
Fu, Chi-An
Lee, Hung-yi
author_facet Liu, Ting-Chun
Hsu, Ching-Yu
Lee, Kuan-Yi
Fu, Chi-An
Lee, Hung-yi
contents Prompt injection attacks pose a significant challenge to the safe deployment of Large Language Models (LLMs) in real-world applications. While prompt-based detection offers a lightweight and interpretable defense strategy, its effectiveness has been hindered by the need for manual prompt engineering. To address this issue, we propose AEGIS , an Automated co-Evolutionary framework for Guarding prompt Injections Schema. Both attack and defense prompts are iteratively optimized against each other using a gradient-like natural language prompt optimization technique. This framework enables both attackers and defenders to autonomously evolve via a Textual Gradient Optimization (TGO) module, leveraging feedback from an LLM-guided evaluation loop. We evaluate our system on a real-world assignment grading dataset of prompt injection attacks and demonstrate that our method consistently outperforms existing baselines, achieving superior robustness in both attack success and detection. Specifically, the attack success rate (ASR) reaches 1.0, representing an improvement of 0.26 over the baseline. For detection, the true positive rate (TPR) improves by 0.23 compared to the previous best work, reaching 0.84, and the true negative rate (TNR) remains comparable at 0.89. Ablation studies confirm the importance of co-evolution, gradient buffering, and multi-objective optimization. We also confirm that this framework is effective in different LLMs. Our results highlight the promise of adversarial training as a scalable and effective approach for guarding prompt injections.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00088
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema
Liu, Ting-Chun
Hsu, Ching-Yu
Lee, Kuan-Yi
Fu, Chi-An
Lee, Hung-yi
Cryptography and Security
Artificial Intelligence
Machine Learning
Prompt injection attacks pose a significant challenge to the safe deployment of Large Language Models (LLMs) in real-world applications. While prompt-based detection offers a lightweight and interpretable defense strategy, its effectiveness has been hindered by the need for manual prompt engineering. To address this issue, we propose AEGIS , an Automated co-Evolutionary framework for Guarding prompt Injections Schema. Both attack and defense prompts are iteratively optimized against each other using a gradient-like natural language prompt optimization technique. This framework enables both attackers and defenders to autonomously evolve via a Textual Gradient Optimization (TGO) module, leveraging feedback from an LLM-guided evaluation loop. We evaluate our system on a real-world assignment grading dataset of prompt injection attacks and demonstrate that our method consistently outperforms existing baselines, achieving superior robustness in both attack success and detection. Specifically, the attack success rate (ASR) reaches 1.0, representing an improvement of 0.26 over the baseline. For detection, the true positive rate (TPR) improves by 0.23 compared to the previous best work, reaching 0.84, and the true negative rate (TNR) remains comparable at 0.89. Ablation studies confirm the importance of co-evolution, gradient buffering, and multi-objective optimization. We also confirm that this framework is effective in different LLMs. Our results highlight the promise of adversarial training as a scalable and effective approach for guarding prompt injections.
title AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.00088