Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hill, Brennen, Parla, Surendra, Balabhadruni, Venkata Abhijeeth, Padmalayam, Atharv Prajod, Sharma, Sujay Chandra Shekara
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911138808070144
author Hill, Brennen
Parla, Surendra
Balabhadruni, Venkata Abhijeeth
Padmalayam, Atharv Prajod
Sharma, Sujay Chandra Shekara
author_facet Hill, Brennen
Parla, Surendra
Balabhadruni, Venkata Abhijeeth
Padmalayam, Atharv Prajod
Sharma, Sujay Chandra Shekara
contents The proliferation of Large Language Models (LLMs) has introduced critical security challenges, where adversarial actors can manipulate input prompts to cause significant harm and circumvent safety alignments. These prompt-based attacks exploit vulnerabilities in a model's design, training, and contextual understanding, leading to intellectual property theft, misinformation generation, and erosion of user trust. A systematic understanding of these attack vectors is the foundational step toward developing robust countermeasures. This paper presents a comprehensive literature survey of prompt-based attack methodologies, categorizing them to provide a clear threat model. By detailing the mechanisms and impacts of these exploits, this survey aims to inform the research community's efforts in building the next generation of secure LLMs that are inherently resistant to unauthorized distillation, fine-tuning, and editing.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
Hill, Brennen
Parla, Surendra
Balabhadruni, Venkata Abhijeeth
Padmalayam, Atharv Prajod
Sharma, Sujay Chandra Shekara
Computation and Language
Cryptography and Security
Machine Learning
68T07, 68T50
I.2.7; I.2.6; K.6.5
The proliferation of Large Language Models (LLMs) has introduced critical security challenges, where adversarial actors can manipulate input prompts to cause significant harm and circumvent safety alignments. These prompt-based attacks exploit vulnerabilities in a model's design, training, and contextual understanding, leading to intellectual property theft, misinformation generation, and erosion of user trust. A systematic understanding of these attack vectors is the foundational step toward developing robust countermeasures. This paper presents a comprehensive literature survey of prompt-based attack methodologies, categorizing them to provide a clear threat model. By detailing the mechanisms and impacts of these exploits, this survey aims to inform the research community's efforts in building the next generation of secure LLMs that are inherently resistant to unauthorized distillation, fine-tuning, and editing.
title Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
topic Computation and Language
Cryptography and Security
Machine Learning
68T07, 68T50
I.2.7; I.2.6; K.6.5
url https://arxiv.org/abs/2509.04615