SoK: Prompt Hacking of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rababah, Baha, Shang, Wu, Kwiatkowski, Matthew, Leung, Carson, Akcora, Cuneyt Gurcan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929549413974016
author Rababah, Baha
Shang
Wu
Kwiatkowski, Matthew
Leung, Carson
Akcora, Cuneyt Gurcan
author_facet Rababah, Baha
Shang
Wu
Kwiatkowski, Matthew
Leung, Carson
Akcora, Cuneyt Gurcan
contents The safety and robustness of large language models (LLMs) based applications remain critical challenges in artificial intelligence. Among the key threats to these applications are prompt hacking attacks, which can significantly undermine the security and reliability of LLM-based systems. In this work, we offer a comprehensive and systematic overview of three distinct types of prompt hacking: jailbreaking, leaking, and injection, addressing the nuances that differentiate them despite their overlapping characteristics. To enhance the evaluation of LLM-based applications, we propose a novel framework that categorizes LLM responses into five distinct classes, moving beyond the traditional binary classification. This approach provides more granular insights into the AI's behavior, improving diagnostic precision and enabling more targeted enhancements to the system's safety and robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13901
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SoK: Prompt Hacking of Large Language Models
Rababah, Baha
Shang
Wu
Kwiatkowski, Matthew
Leung, Carson
Akcora, Cuneyt Gurcan
Cryptography and Security
Artificial Intelligence
Computation and Language
Emerging Technologies
The safety and robustness of large language models (LLMs) based applications remain critical challenges in artificial intelligence. Among the key threats to these applications are prompt hacking attacks, which can significantly undermine the security and reliability of LLM-based systems. In this work, we offer a comprehensive and systematic overview of three distinct types of prompt hacking: jailbreaking, leaking, and injection, addressing the nuances that differentiate them despite their overlapping characteristics. To enhance the evaluation of LLM-based applications, we propose a novel framework that categorizes LLM responses into five distinct classes, moving beyond the traditional binary classification. This approach provides more granular insights into the AI's behavior, improving diagnostic precision and enabling more targeted enhancements to the system's safety and robustness.
title SoK: Prompt Hacking of Large Language Models
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Emerging Technologies
url https://arxiv.org/abs/2410.13901