Efficient Universal Goal Hijacking with Semantics-guided Prompt Organization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yihao, Wang, Chong, Jia, Xiaojun, Guo, Qing, Juefei-Xu, Felix, Zhang, Jian, Pu, Geguang, Liu, Yang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910975318294528
author Huang, Yihao
Wang, Chong
Jia, Xiaojun
Guo, Qing
Juefei-Xu, Felix
Zhang, Jian
Pu, Geguang
Liu, Yang
author_facet Huang, Yihao
Wang, Chong
Jia, Xiaojun
Guo, Qing
Juefei-Xu, Felix
Zhang, Jian
Pu, Geguang
Liu, Yang
contents Universal goal hijacking is a kind of prompt injection attack that forces LLMs to return a target malicious response for arbitrary normal user prompts. The previous methods achieve high attack performance while being too cumbersome and time-consuming. Also, they have concentrated solely on optimization algorithms, overlooking the crucial role of the prompt. To this end, we propose a method called POUGH that incorporates an efficient optimization algorithm and two semantics-guided prompt organization strategies. Specifically, our method starts with a sampling strategy to select representative prompts from a candidate pool, followed by a ranking strategy that prioritizes them. Given the sequentially ranked prompts, our method employs an iterative optimization algorithm to generate a fixed suffix that can concatenate to arbitrary user prompts for universal goal hijacking. Experiments conducted on four popular LLMs and ten types of target responses verified the effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14189
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Universal Goal Hijacking with Semantics-guided Prompt Organization
Huang, Yihao
Wang, Chong
Jia, Xiaojun
Guo, Qing
Juefei-Xu, Felix
Zhang, Jian
Pu, Geguang
Liu, Yang
Computation and Language
Computer Vision and Pattern Recognition
Universal goal hijacking is a kind of prompt injection attack that forces LLMs to return a target malicious response for arbitrary normal user prompts. The previous methods achieve high attack performance while being too cumbersome and time-consuming. Also, they have concentrated solely on optimization algorithms, overlooking the crucial role of the prompt. To this end, we propose a method called POUGH that incorporates an efficient optimization algorithm and two semantics-guided prompt organization strategies. Specifically, our method starts with a sampling strategy to select representative prompts from a candidate pool, followed by a ranking strategy that prioritizes them. Given the sequentially ranked prompts, our method employs an iterative optimization algorithm to generate a fixed suffix that can concatenate to arbitrary user prompts for universal goal hijacking. Experiments conducted on four popular LLMs and ten types of target responses verified the effectiveness.
title Efficient Universal Goal Hijacking with Semantics-guided Prompt Organization
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.14189