System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Fangzhou, Cecchetti, Ethan, Xiao, Chaowei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913539836346368
author Wu, Fangzhou
Cecchetti, Ethan
Xiao, Chaowei
author_facet Wu, Fangzhou
Cecchetti, Ethan
Xiao, Chaowei
contents Large Language Model-based systems (LLM systems) are information and query processing systems that use LLMs to plan operations from natural-language prompts and feed the output of each successive step into the LLM to plan the next. This structure results in powerful tools that can process complex information from diverse sources but raises critical security concerns. Malicious information from any source may be processed by the LLM and can compromise the query processing, resulting in nearly arbitrary misbehavior. To tackle this problem, we present a system-level defense based on the principles of information flow control that we call an f-secure LLM system. An f-secure LLM system disaggregates the components of an LLM system into a context-aware pipeline with dynamically generated structured executable plans, and a security monitor filters out untrusted input into the planning process. This structure prevents compromise while maximizing flexibility. We provide formal models for both existing LLM systems and our f-secure LLM system, allowing analysis of critical security guarantees. We further evaluate case studies and benchmarks showing that f-secure LLM systems provide robust security while preserving functionality and efficiency. Our code is released at https://github.com/fzwark/Secure_LLM_System.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19091
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
Wu, Fangzhou
Cecchetti, Ethan
Xiao, Chaowei
Cryptography and Security
Large Language Model-based systems (LLM systems) are information and query processing systems that use LLMs to plan operations from natural-language prompts and feed the output of each successive step into the LLM to plan the next. This structure results in powerful tools that can process complex information from diverse sources but raises critical security concerns. Malicious information from any source may be processed by the LLM and can compromise the query processing, resulting in nearly arbitrary misbehavior. To tackle this problem, we present a system-level defense based on the principles of information flow control that we call an f-secure LLM system. An f-secure LLM system disaggregates the components of an LLM system into a context-aware pipeline with dynamically generated structured executable plans, and a security monitor filters out untrusted input into the planning process. This structure prevents compromise while maximizing flexibility. We provide formal models for both existing LLM systems and our f-secure LLM system, allowing analysis of critical security guarantees. We further evaluate case studies and benchmarks showing that f-secure LLM systems provide robust security while preserving functionality and efficiency. Our code is released at https://github.com/fzwark/Secure_LLM_System.
title System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
topic Cryptography and Security
url https://arxiv.org/abs/2409.19091