Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Momcilovic, Tomas Bueno, Balta, Dian, Buesser, Beat, Zizzo, Giulio, Purcell, Mark
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913535875874816
author Momcilovic, Tomas Bueno
Balta, Dian
Buesser, Beat
Zizzo, Giulio
Purcell, Mark
author_facet Momcilovic, Tomas Bueno
Balta, Dian
Buesser, Beat
Zizzo, Giulio
Purcell, Mark
contents This paper presents an approach to developing assurance cases for adversarial robustness and regulatory compliance in large language models (LLMs). Focusing on both natural and code language tasks, we explore the vulnerabilities these models face, including adversarial attacks based on jailbreaking, heuristics, and randomization. We propose a layered framework incorporating guardrails at various stages of LLM deployment, aimed at mitigating these attacks and ensuring compliance with the EU AI Act. Our approach includes a meta-layer for dynamic risk management and reasoning, crucial for addressing the evolving nature of LLM vulnerabilities. We illustrate our method with two exemplary assurance cases, highlighting how different contexts demand tailored strategies to ensure robust and compliant AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05304
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
Momcilovic, Tomas Bueno
Balta, Dian
Buesser, Beat
Zizzo, Giulio
Purcell, Mark
Cryptography and Security
Artificial Intelligence
Software Engineering
This paper presents an approach to developing assurance cases for adversarial robustness and regulatory compliance in large language models (LLMs). Focusing on both natural and code language tasks, we explore the vulnerabilities these models face, including adversarial attacks based on jailbreaking, heuristics, and randomization. We propose a layered framework incorporating guardrails at various stages of LLM deployment, aimed at mitigating these attacks and ensuring compliance with the EU AI Act. Our approach includes a meta-layer for dynamic risk management and reasoning, crucial for addressing the evolving nature of LLM vulnerabilities. We illustrate our method with two exemplary assurance cases, highlighting how different contexts demand tailored strategies to ensure robust and compliant AI systems.
title Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
topic Cryptography and Security
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2410.05304