SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Hanbin, Feng, Shuya, Naderloui, Nima, Yan, Shenao, Zhang, Jingyu, Liu, Biying, Arastehfard, Ali, Huang, Heqing, Hong, Yuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914104848941056
author Hong, Hanbin
Feng, Shuya
Naderloui, Nima
Yan, Shenao
Zhang, Jingyu
Liu, Biying
Arastehfard, Ali
Huang, Heqing
Hong, Yuan
author_facet Hong, Hanbin
Feng, Shuya
Naderloui, Nima
Yan, Shenao
Zhang, Jingyu
Liu, Biying
Arastehfard, Ali
Huang, Heqing
Hong, Yuan
contents Large Language Models (LLMs) have rapidly become integral to real-world applications, powering services across diverse sectors. However, their widespread deployment has exposed critical security risks, particularly through jailbreak prompts that can bypass model alignment and induce harmful outputs. Despite intense research into both attack and defense techniques, the field remains fragmented: definitions, threat models, and evaluation criteria vary widely, impeding systematic progress and fair comparison. In this Systematization of Knowledge (SoK), we address these challenges by (1) proposing a holistic, multi-level taxonomy that organizes attacks, defenses, and vulnerabilities in LLM prompt security; (2) formalizing threat models and cost assumptions into machine-readable profiles for reproducible evaluation; (3) introducing an open-source evaluation toolkit for standardized, auditable comparison of attacks and defenses; (4) releasing JAILBREAKDB, the largest annotated dataset of jailbreak and benign prompts to date;\footnote{The dataset is released at \href{https://huggingface.co/datasets/youbin2014/JailbreakDB}{\textcolor{purple}{https://huggingface.co/datasets/youbin2014/JailbreakDB}}.} and (5) presenting a comprehensive evaluation platform and leaderboard of state-of-the-art methods \footnote{will be released soon.}. Our work unifies fragmented research, provides rigorous foundations for future studies, and supports the development of robust, trustworthy LLMs suitable for high-stakes deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
Hong, Hanbin
Feng, Shuya
Naderloui, Nima
Yan, Shenao
Zhang, Jingyu
Liu, Biying
Arastehfard, Ali
Huang, Heqing
Hong, Yuan
Cryptography and Security
Artificial Intelligence
Large Language Models (LLMs) have rapidly become integral to real-world applications, powering services across diverse sectors. However, their widespread deployment has exposed critical security risks, particularly through jailbreak prompts that can bypass model alignment and induce harmful outputs. Despite intense research into both attack and defense techniques, the field remains fragmented: definitions, threat models, and evaluation criteria vary widely, impeding systematic progress and fair comparison. In this Systematization of Knowledge (SoK), we address these challenges by (1) proposing a holistic, multi-level taxonomy that organizes attacks, defenses, and vulnerabilities in LLM prompt security; (2) formalizing threat models and cost assumptions into machine-readable profiles for reproducible evaluation; (3) introducing an open-source evaluation toolkit for standardized, auditable comparison of attacks and defenses; (4) releasing JAILBREAKDB, the largest annotated dataset of jailbreak and benign prompts to date;\footnote{The dataset is released at \href{https://huggingface.co/datasets/youbin2014/JailbreakDB}{\textcolor{purple}{https://huggingface.co/datasets/youbin2014/JailbreakDB}}.} and (5) presenting a comprehensive evaluation platform and leaderboard of state-of-the-art methods \footnote{will be released soon.}. Our work unifies fragmented research, provides rigorous foundations for future studies, and supports the development of robust, trustworthy LLMs suitable for high-stakes deployment.
title SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.15476