MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Muyang, Yao, Yuanzhi, Lin, Changting, Kai, Caihong, Chen, Yanxiang, Liu, Zhiquan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916958674354176
author Zheng, Muyang
Yao, Yuanzhi
Lin, Changting
Kai, Caihong
Chen, Yanxiang
Liu, Zhiquan
author_facet Zheng, Muyang
Yao, Yuanzhi
Lin, Changting
Kai, Caihong
Chen, Yanxiang
Liu, Zhiquan
contents Despite efforts to align large language models (LLMs) with societal and moral values, these models remain susceptible to jailbreak attacks -- methods designed to elicit harmful responses. Jailbreaking black-box LLMs is considered challenging due to the discrete nature of token inputs, restricted access to the target LLM, and limited query budget. To address the issues above, we propose an effective method for jailbreaking black-box large language Models via Iterative Semantic Tuning, named MIST. MIST enables attackers to iteratively refine prompts that preserve the original semantic intent while inducing harmful content. Specifically, to balance semantic similarity with computational efficiency, MIST incorporates two key strategies: sequential synonym search, and its advanced version -- order-determining optimization. We conduct extensive experiments on two datasets using two open-source and four closed-source models. Results show that MIST achieves competitive attack success rate, relatively low query count, and fair transferability, outperforming or matching state-of-the-art jailbreak methods. Additionally, we conduct analysis on computational efficiency to validate the practical viability of MIST.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16792
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning
Zheng, Muyang
Yao, Yuanzhi
Lin, Changting
Kai, Caihong
Chen, Yanxiang
Liu, Zhiquan
Computation and Language
Artificial Intelligence
Despite efforts to align large language models (LLMs) with societal and moral values, these models remain susceptible to jailbreak attacks -- methods designed to elicit harmful responses. Jailbreaking black-box LLMs is considered challenging due to the discrete nature of token inputs, restricted access to the target LLM, and limited query budget. To address the issues above, we propose an effective method for jailbreaking black-box large language Models via Iterative Semantic Tuning, named MIST. MIST enables attackers to iteratively refine prompts that preserve the original semantic intent while inducing harmful content. Specifically, to balance semantic similarity with computational efficiency, MIST incorporates two key strategies: sequential synonym search, and its advanced version -- order-determining optimization. We conduct extensive experiments on two datasets using two open-source and four closed-source models. Results show that MIST achieves competitive attack success rate, relatively low query count, and fair transferability, outperforming or matching state-of-the-art jailbreak methods. Additionally, we conduct analysis on computational efficiency to validate the practical viability of MIST.
title MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.16792