Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yeo, Andrew, Choi, Daeseon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912574745870336
author Yeo, Andrew
Choi, Daeseon
author_facet Yeo, Andrew
Choi, Daeseon
contents Large Language Models (LLMs) have seen rapid adoption in recent years, with industries increasingly relying on them to maintain a competitive advantage. These models excel at interpreting user instructions and generating human-like responses, leading to their integration across diverse domains, including consulting and information retrieval. However, their widespread deployment also introduces substantial security risks, most notably in the form of prompt injection and jailbreak attacks. To systematically evaluate LLM vulnerabilities -- particularly to external prompt injection -- we conducted a series of experiments on eight commercial models. Each model was tested without supplementary sanitization, relying solely on its built-in safeguards. The results exposed exploitable weaknesses and emphasized the need for stronger security measures. Four categories of attacks were examined: direct injection, indirect (external) injection, image-based injection, and prompt leakage. Comparative analysis indicated that Claude 3 demonstrated relatively greater robustness; nevertheless, empirical findings confirm that additional defenses, such as input normalization, remain necessary to achieve reliable protection.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05883
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
Yeo, Andrew
Choi, Daeseon
Cryptography and Security
Artificial Intelligence
Large Language Models (LLMs) have seen rapid adoption in recent years, with industries increasingly relying on them to maintain a competitive advantage. These models excel at interpreting user instructions and generating human-like responses, leading to their integration across diverse domains, including consulting and information retrieval. However, their widespread deployment also introduces substantial security risks, most notably in the form of prompt injection and jailbreak attacks. To systematically evaluate LLM vulnerabilities -- particularly to external prompt injection -- we conducted a series of experiments on eight commercial models. Each model was tested without supplementary sanitization, relying solely on its built-in safeguards. The results exposed exploitable weaknesses and emphasized the need for stronger security measures. Four categories of attacks were examined: direct injection, indirect (external) injection, image-based injection, and prompt leakage. Comparative analysis indicated that Claude 3 demonstrated relatively greater robustness; nevertheless, empirical findings confirm that additional defenses, such as input normalization, remain necessary to achieve reliable protection.
title Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2509.05883