Prompt Injection Attacks on Large Language Models in Oncology

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Clusmann, Jan, Ferber, Dyke, Wiest, Isabella C., Schneider, Carolin V., Brinker, Titus J., Foersch, Sebastian, Truhn, Daniel, Kather, Jakob N.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908273839439872
author Clusmann, Jan
Ferber, Dyke
Wiest, Isabella C.
Schneider, Carolin V.
Brinker, Titus J.
Foersch, Sebastian
Truhn, Daniel
Kather, Jakob N.
author_facet Clusmann, Jan
Ferber, Dyke
Wiest, Isabella C.
Schneider, Carolin V.
Brinker, Titus J.
Foersch, Sebastian
Truhn, Daniel
Kather, Jakob N.
contents Vision-language artificial intelligence models (VLMs) possess medical knowledge and can be employed in healthcare in numerous ways, including as image interpreters, virtual scribes, and general decision support systems. However, here, we demonstrate that current VLMs applied to medical tasks exhibit a fundamental security flaw: they can be attacked by prompt injection attacks, which can be used to output harmful information just by interacting with the VLM, without any access to its parameters. We performed a quantitative study to evaluate the vulnerabilities to these attacks in four state of the art VLMs which have been proposed to be of utility in healthcare: Claude 3 Opus, Claude 3.5 Sonnet, Reka Core, and GPT-4o. Using a set of N=297 attacks, we show that all of these models are susceptible. Specifically, we show that embedding sub-visual prompts in medical imaging data can cause the model to provide harmful output, and that these prompts are non-obvious to human observers. Thus, our study demonstrates a key vulnerability in medical VLMs which should be mitigated before widespread clinical adoption.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18981
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prompt Injection Attacks on Large Language Models in Oncology
Clusmann, Jan
Ferber, Dyke
Wiest, Isabella C.
Schneider, Carolin V.
Brinker, Titus J.
Foersch, Sebastian
Truhn, Daniel
Kather, Jakob N.
Cryptography and Security
Artificial Intelligence
Machine Learning
Vision-language artificial intelligence models (VLMs) possess medical knowledge and can be employed in healthcare in numerous ways, including as image interpreters, virtual scribes, and general decision support systems. However, here, we demonstrate that current VLMs applied to medical tasks exhibit a fundamental security flaw: they can be attacked by prompt injection attacks, which can be used to output harmful information just by interacting with the VLM, without any access to its parameters. We performed a quantitative study to evaluate the vulnerabilities to these attacks in four state of the art VLMs which have been proposed to be of utility in healthcare: Claude 3 Opus, Claude 3.5 Sonnet, Reka Core, and GPT-4o. Using a set of N=297 attacks, we show that all of these models are susceptible. Specifically, we show that embedding sub-visual prompts in medical imaging data can cause the model to provide harmful output, and that these prompts are non-obvious to human observers. Thus, our study demonstrates a key vulnerability in medical VLMs which should be mitigated before widespread clinical adoption.
title Prompt Injection Attacks on Large Language Models in Oncology
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.18981