Defense against Prompt Injection Attacks via Mixture of Encodings

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Ruiyi, Sullivan, David, Jackson, Kyle, Xie, Pengtao, Chen, Mei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915235831480320
author Zhang, Ruiyi
Sullivan, David
Jackson, Kyle
Xie, Pengtao
Chen, Mei
author_facet Zhang, Ruiyi
Sullivan, David
Jackson, Kyle
Xie, Pengtao
Chen, Mei
contents Large Language Models (LLMs) have emerged as a dominant approach for a wide range of NLP tasks, with their access to external information further enhancing their capabilities. However, this introduces new vulnerabilities, known as prompt injection attacks, where external content embeds malicious instructions that manipulate the LLM's output. Recently, the Base64 defense has been recognized as one of the most effective methods for reducing success rate of prompt injection attacks. Despite its efficacy, this method can degrade LLM performance on certain NLP tasks. To address this challenge, we propose a novel defense mechanism: mixture of encodings, which utilizes multiple character encodings, including Base64. Extensive experimental results show that our method achieves one of the lowest attack success rates under prompt injection attacks, while maintaining high performance across all NLP tasks, outperforming existing character encoding-based defense methods. This underscores the effectiveness of our mixture of encodings strategy for both safety and task performance metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2504_07467
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Defense against Prompt Injection Attacks via Mixture of Encodings
Zhang, Ruiyi
Sullivan, David
Jackson, Kyle
Xie, Pengtao
Chen, Mei
Computation and Language
Large Language Models (LLMs) have emerged as a dominant approach for a wide range of NLP tasks, with their access to external information further enhancing their capabilities. However, this introduces new vulnerabilities, known as prompt injection attacks, where external content embeds malicious instructions that manipulate the LLM's output. Recently, the Base64 defense has been recognized as one of the most effective methods for reducing success rate of prompt injection attacks. Despite its efficacy, this method can degrade LLM performance on certain NLP tasks. To address this challenge, we propose a novel defense mechanism: mixture of encodings, which utilizes multiple character encodings, including Base64. Extensive experimental results show that our method achieves one of the lowest attack success rates under prompt injection attacks, while maintaining high performance across all NLP tasks, outperforming existing character encoding-based defense methods. This underscores the effectiveness of our mixture of encodings strategy for both safety and task performance metrics.
title Defense against Prompt Injection Attacks via Mixture of Encodings
topic Computation and Language
url https://arxiv.org/abs/2504.07467