Saved in:
Bibliographic Details
Main Authors: Lan, Yifan, Cao, Yuanpu, Zhang, Weitong, Lin, Lu, Chen, Jinghui
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.12521
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912588811468800
author Lan, Yifan
Cao, Yuanpu
Zhang, Weitong
Lin, Lu
Chen, Jinghui
author_facet Lan, Yifan
Cao, Yuanpu
Zhang, Weitong
Lin, Lu
Chen, Jinghui
contents Recently, Multimodal Large Language Models (MLLMs) have gained significant attention across various domains. However, their widespread adoption has also raised serious safety concerns. In this paper, we uncover a new safety risk of MLLMs: the output preference of MLLMs can be arbitrarily manipulated by carefully optimized images. Such attacks often generate contextually relevant yet biased responses that are neither overtly harmful nor unethical, making them difficult to detect. Specifically, we introduce a novel method, Preference Hijacking (Phi), for manipulating the MLLM response preferences using a preference hijacked image. Our method works at inference time and requires no model modifications. Additionally, we introduce a universal hijacking perturbation -- a transferable component that can be embedded into different images to hijack MLLM responses toward any attacker-specified preferences. Experimental results across various tasks demonstrate the effectiveness of our approach. The code for Phi is accessible at https://github.com/Yifan-Lan/Phi.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12521
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
Lan, Yifan
Cao, Yuanpu
Zhang, Weitong
Lin, Lu
Chen, Jinghui
Machine Learning
Recently, Multimodal Large Language Models (MLLMs) have gained significant attention across various domains. However, their widespread adoption has also raised serious safety concerns. In this paper, we uncover a new safety risk of MLLMs: the output preference of MLLMs can be arbitrarily manipulated by carefully optimized images. Such attacks often generate contextually relevant yet biased responses that are neither overtly harmful nor unethical, making them difficult to detect. Specifically, we introduce a novel method, Preference Hijacking (Phi), for manipulating the MLLM response preferences using a preference hijacked image. Our method works at inference time and requires no model modifications. Additionally, we introduce a universal hijacking perturbation -- a transferable component that can be embedded into different images to hijack MLLM responses toward any attacker-specified preferences. Experimental results across various tasks demonstrate the effectiveness of our approach. The code for Phi is accessible at https://github.com/Yifan-Lan/Phi.
title Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
topic Machine Learning
url https://arxiv.org/abs/2509.12521