Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915355990949888 |
|---|---|
| author | Zhang, Chong Li, Xiang Wang, Jia Liang, Shan Xue, Haochen Jin, Xiaobo |
| author_facet | Zhang, Chong Li, Xiang Wang, Jia Liang, Shan Xue, Haochen Jin, Xiaobo |
| contents | Large Language Models (LLMs) increasingly rely on automatic prompt engineering in graphical user interfaces (GUIs) to refine user inputs and enhance response accuracy. However, the diversity of user requirements often leads to unintended misinterpretations, where automated optimizations distort original intentions and produce erroneous outputs. To address this challenge, we propose the Adaptive Greedy Binary Search (AGBS) method, which simulates common prompt optimization mechanisms while preserving semantic stability. Our approach dynamically evaluates the impact of such strategies on LLM performance, enabling robust adversarial sample generation. Through extensive experiments on open and closed-source LLMs, we demonstrate AGBS's effectiveness in balancing semantic consistency and attack efficacy. Our findings offer actionable insights for designing more reliable prompt optimization systems. Code is available at: https://github.com/franz-chang/DOBS |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18756 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach Zhang, Chong Li, Xiang Wang, Jia Liang, Shan Xue, Haochen Jin, Xiaobo Computation and Language Cryptography and Security Large Language Models (LLMs) increasingly rely on automatic prompt engineering in graphical user interfaces (GUIs) to refine user inputs and enhance response accuracy. However, the diversity of user requirements often leads to unintended misinterpretations, where automated optimizations distort original intentions and produce erroneous outputs. To address this challenge, we propose the Adaptive Greedy Binary Search (AGBS) method, which simulates common prompt optimization mechanisms while preserving semantic stability. Our approach dynamically evaluates the impact of such strategies on LLM performance, enabling robust adversarial sample generation. Through extensive experiments on open and closed-source LLMs, we demonstrate AGBS's effectiveness in balancing semantic consistency and attack efficacy. Our findings offer actionable insights for designing more reliable prompt optimization systems. Code is available at: https://github.com/franz-chang/DOBS |
| title | Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach |
| topic | Computation and Language Cryptography and Security |
| url | https://arxiv.org/abs/2506.18756 |