Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chong, Li, Xiang, Wang, Jia, Liang, Shan, Xue, Haochen, Jin, Xiaobo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915355990949888
author Zhang, Chong
Li, Xiang
Wang, Jia
Liang, Shan
Xue, Haochen
Jin, Xiaobo
author_facet Zhang, Chong
Li, Xiang
Wang, Jia
Liang, Shan
Xue, Haochen
Jin, Xiaobo
contents Large Language Models (LLMs) increasingly rely on automatic prompt engineering in graphical user interfaces (GUIs) to refine user inputs and enhance response accuracy. However, the diversity of user requirements often leads to unintended misinterpretations, where automated optimizations distort original intentions and produce erroneous outputs. To address this challenge, we propose the Adaptive Greedy Binary Search (AGBS) method, which simulates common prompt optimization mechanisms while preserving semantic stability. Our approach dynamically evaluates the impact of such strategies on LLM performance, enabling robust adversarial sample generation. Through extensive experiments on open and closed-source LLMs, we demonstrate AGBS's effectiveness in balancing semantic consistency and attack efficacy. Our findings offer actionable insights for designing more reliable prompt optimization systems. Code is available at: https://github.com/franz-chang/DOBS
format Preprint
id arxiv_https___arxiv_org_abs_2506_18756
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
Zhang, Chong
Li, Xiang
Wang, Jia
Liang, Shan
Xue, Haochen
Jin, Xiaobo
Computation and Language
Cryptography and Security
Large Language Models (LLMs) increasingly rely on automatic prompt engineering in graphical user interfaces (GUIs) to refine user inputs and enhance response accuracy. However, the diversity of user requirements often leads to unintended misinterpretations, where automated optimizations distort original intentions and produce erroneous outputs. To address this challenge, we propose the Adaptive Greedy Binary Search (AGBS) method, which simulates common prompt optimization mechanisms while preserving semantic stability. Our approach dynamically evaluates the impact of such strategies on LLM performance, enabling robust adversarial sample generation. Through extensive experiments on open and closed-source LLMs, we demonstrate AGBS's effectiveness in balancing semantic consistency and attack efficacy. Our findings offer actionable insights for designing more reliable prompt optimization systems. Code is available at: https://github.com/franz-chang/DOBS
title Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
topic Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2506.18756