A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Swanda, Adam, Chang, Amy, Chen, Alexander, Burch, Fraser, Kassianik, Paul, Berlin, Konstantin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915559669497856
author Swanda, Adam
Chang, Amy
Chen, Alexander
Burch, Fraser
Kassianik, Paul
Berlin, Konstantin
author_facet Swanda, Adam
Chang, Amy
Chen, Alexander
Burch, Fraser
Kassianik, Paul
Berlin, Konstantin
contents The widespread adoption of Large Language Models (LLMs) has revolutionized AI deployment, enabling autonomous and semi-autonomous applications across industries through intuitive language interfaces and continuous improvements in model development. However, the attendant increase in autonomy and expansion of access permissions among AI applications also make these systems compelling targets for malicious attacks. Their inherent susceptibility to security flaws necessitates robust defenses, yet no known approaches can prevent zero-day or novel attacks against LLMs. This places AI protection systems in a category similar to established malware protection systems: rather than providing guaranteed immunity, they minimize risk through enhanced observability, multi-layered defense, and rapid threat response, supported by a threat intelligence function designed specifically for AI-related threats. Prior work on LLM protection has largely evaluated individual detection models rather than end-to-end systems designed for continuous, rapid adaptation to a changing threat landscape. We present a production-grade defense system rooted in established malware detection and threat intelligence practices. Our platform integrates three components: a threat intelligence system that turns emerging threats into protections; a data platform that aggregates and enriches information while providing observability, monitoring, and ML operations; and a release platform enabling safe, rapid detection updates without disrupting customer workflows. Together, these components deliver layered protection against evolving LLM threats while generating training data for continuous model improvement and deploying updates without interrupting production.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks
Swanda, Adam
Chang, Amy
Chen, Alexander
Burch, Fraser
Kassianik, Paul
Berlin, Konstantin
Cryptography and Security
Artificial Intelligence
The widespread adoption of Large Language Models (LLMs) has revolutionized AI deployment, enabling autonomous and semi-autonomous applications across industries through intuitive language interfaces and continuous improvements in model development. However, the attendant increase in autonomy and expansion of access permissions among AI applications also make these systems compelling targets for malicious attacks. Their inherent susceptibility to security flaws necessitates robust defenses, yet no known approaches can prevent zero-day or novel attacks against LLMs. This places AI protection systems in a category similar to established malware protection systems: rather than providing guaranteed immunity, they minimize risk through enhanced observability, multi-layered defense, and rapid threat response, supported by a threat intelligence function designed specifically for AI-related threats. Prior work on LLM protection has largely evaluated individual detection models rather than end-to-end systems designed for continuous, rapid adaptation to a changing threat landscape. We present a production-grade defense system rooted in established malware detection and threat intelligence practices. Our platform integrates three components: a threat intelligence system that turns emerging threats into protections; a data platform that aggregates and enriches information while providing observability, monitoring, and ML operations; and a release platform enabling safe, rapid detection updates without disrupting customer workflows. Together, these components deliver layered protection against evolving LLM threats while generating training data for continuous model improvement and deploying updates without interrupting production.
title A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2509.20639