Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection
Fuente:
arXiv
Saved in:
| Main Authors: | Shao, Zedian, Liu, Hongbin, Hu, Yuepeng, Gong, Neil Zhenqiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Refusing Safe Prompts for Multi-modal Large Language Models
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
A Critical Evaluation of Defenses against Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2025)
by: Jia, Yuqi, et al.
Published: (2025)
Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
by: Jiang, Zhengyuan, et al.
Published: (2025)
by: Jiang, Zhengyuan, et al.
Published: (2025)
SafeText: Safe Text-to-image Models via Aligning the Text Encoder
by: Hu, Yuepeng, et al.
Published: (2025)
by: Hu, Yuepeng, et al.
Published: (2025)
PromptLocate: Localizing Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2025)
by: Jia, Yuqi, et al.
Published: (2025)
Certifiably Robust Image Watermark
by: Jiang, Zhengyuan, et al.
Published: (2024)
by: Jiang, Zhengyuan, et al.
Published: (2024)
ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
by: Wang, Reachal, et al.
Published: (2025)
by: Wang, Reachal, et al.
Published: (2025)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
by: Liu, Yupei, et al.
Published: (2025)
by: Liu, Yupei, et al.
Published: (2025)
A Transfer Attack to Image Watermarks
by: Hu, Yuepeng, et al.
Published: (2024)
by: Hu, Yuepeng, et al.
Published: (2024)
Automatically Generating Visual Hallucination Test Cases for Multimodal Large Language Models
by: Liu, Zhongye, et al.
Published: (2024)
by: Liu, Zhongye, et al.
Published: (2024)
Stable Signature is Unstable: Removing Image Watermark from Diffusion Models
by: Hu, Yuepeng, et al.
Published: (2024)
by: Hu, Yuepeng, et al.
Published: (2024)
Fingerprinting LLMs via Prompt Injection
by: Hu, Yuepeng, et al.
Published: (2025)
by: Hu, Yuepeng, et al.
Published: (2025)
Robustness of Vision Foundation Models to Common Perturbations
by: Liu, Hongbin, et al.
Published: (2026)
by: Liu, Hongbin, et al.
Published: (2026)
Prompt Injection Attack to Tool Selection in LLM Agents
by: Shi, Jiawen, et al.
Published: (2025)
by: Shi, Jiawen, et al.
Published: (2025)
A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Mudjacking: Patching Backdoor Vulnerabilities in Foundation Models
by: Liu, Hongbin, et al.
Published: (2024)
by: Liu, Hongbin, et al.
Published: (2024)
Securing Visually-Aware Recommender Systems: An Adversarial Image Reconstruction and Detection Framework
by: Yin, Minglei, et al.
Published: (2023)
by: Yin, Minglei, et al.
Published: (2023)
Tracing Back the Malicious Clients in Poisoning Attacks to Federated Learning
by: Jia, Yuqi, et al.
Published: (2024)
by: Jia, Yuqi, et al.
Published: (2024)
CorruptEncoder: Data Poisoning based Backdoor Attacks to Contrastive Learning
by: Zhang, Jinghuai, et al.
Published: (2022)
by: Zhang, Jinghuai, et al.
Published: (2022)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
by: Liu, Yupei, et al.
Published: (2025)
by: Liu, Yupei, et al.
Published: (2025)
Model Poisoning Attacks to Federated Learning via Multi-Round Consistency
by: Xie, Yueqi, et al.
Published: (2024)
by: Xie, Yueqi, et al.
Published: (2024)
WebInject: Prompt Injection Attack to Web Agents
by: Wang, Xilong, et al.
Published: (2025)
by: Wang, Xilong, et al.
Published: (2025)
AI-generated Image Detection: Passive or Watermark?
by: Guo, Moyang, et al.
Published: (2024)
by: Guo, Moyang, et al.
Published: (2024)
Watermark-based Attribution of AI-Generated Content
by: Jiang, Zhengyuan, et al.
Published: (2024)
by: Jiang, Zhengyuan, et al.
Published: (2024)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
by: Yuan, Zenghui, et al.
Published: (2025)
by: Yuan, Zenghui, et al.
Published: (2025)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
by: Shi, Jiawen, et al.
Published: (2024)
by: Shi, Jiawen, et al.
Published: (2024)
EditTrack: Detecting and Attributing AI-assisted Image Editing
by: Jiang, Zhengyuan, et al.
Published: (2025)
by: Jiang, Zhengyuan, et al.
Published: (2025)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
by: Liu, Yupei, et al.
Published: (2023)
by: Liu, Yupei, et al.
Published: (2023)
Watermark Robustness and Radioactivity May Be at Odds in Federated Learning
by: Huang, Leixu, et al.
Published: (2025)
by: Huang, Leixu, et al.
Published: (2025)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
by: Nagaraja, Neha, et al.
Published: (2026)
by: Nagaraja, Neha, et al.
Published: (2026)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2026)
by: Jia, Yuqi, et al.
Published: (2026)
VideoMarkBench: Benchmarking Robustness of Video Watermarking
by: Jiang, Zhengyuan, et al.
Published: (2025)
by: Jiang, Zhengyuan, et al.
Published: (2025)
Instance-Level Data-Use Auditing of Visual ML Models
by: Huang, Zonghao, et al.
Published: (2025)
by: Huang, Zonghao, et al.
Published: (2025)
Are You Copying My Prompt? Protecting the Copyright of Vision Prompt for VPaaS via Watermark
by: Ren, Huali, et al.
Published: (2024)
by: Ren, Huali, et al.
Published: (2024)
MalTool: Malicious Tool Attacks on LLM Agents
by: Hu, Yuepeng, et al.
Published: (2026)
by: Hu, Yuepeng, et al.
Published: (2026)
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
by: Tian, Yuan, et al.
Published: (2026)
by: Tian, Yuan, et al.
Published: (2026)
Preventing Prompt Injection with Type-Directed Privilege Separation
by: Jacob, Dennis, et al.
Published: (2025)
by: Jacob, Dennis, et al.
Published: (2025)
Similar Items
-
Refusing Safe Prompts for Multi-modal Large Language Models
by: Shao, Zedian, et al.
Published: (2024) -
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024) -
A Critical Evaluation of Defenses against Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2025) -
Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
by: Jiang, Zhengyuan, et al.
Published: (2025) -
SafeText: Safe Text-to-image Models via Aligning the Text Encoder
by: Hu, Yuepeng, et al.
Published: (2025)