Saved in:
Bibliographic Details
Main Authors: Zhao, Andrew, Ghosh, Reshmi, Carvalho, Vitor, Lawton, Emily, Hines, Keegan, Huang, Gao, Stokes, Jack W.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.14381
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917197893337088
author Zhao, Andrew
Ghosh, Reshmi
Carvalho, Vitor
Lawton, Emily
Hines, Keegan
Huang, Gao
Stokes, Jack W.
author_facet Zhao, Andrew
Ghosh, Reshmi
Carvalho, Vitor
Lawton, Emily
Hines, Keegan
Huang, Gao
Stokes, Jack W.
contents Large language model (LLM) systems increasingly power everyday AI applications such as chatbots, computer-use assistants, and autonomous robots, where performance often depends on manually well-crafted prompts. LLM-based prompt optimizers reduce that effort by iteratively refining prompts from scored feedback, yet the security of this optimization stage remains underexamined. We present the first systematic analysis of poisoning risks in LLM-based prompt optimization. Using HarmBench, we find systems are substantially more vulnerable to manipulated feedback than to query poisoning alone: feedback-based attacks raise attack success rate (ASR) by up to ΔASR = 0.48. We introduce a simple fake reward attack that requires no access to the reward model and significantly increases vulnerability. We also propose a lightweight highlighting defense that reduces the fake reward ΔASR from 0.23 to 0.07 without degrading utility. These results establish prompt optimization pipelines as a first-class attack surface and motivate stronger safeguards for feedback channels and optimization frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14381
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
Zhao, Andrew
Ghosh, Reshmi
Carvalho, Vitor
Lawton, Emily
Hines, Keegan
Huang, Gao
Stokes, Jack W.
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
Large language model (LLM) systems increasingly power everyday AI applications such as chatbots, computer-use assistants, and autonomous robots, where performance often depends on manually well-crafted prompts. LLM-based prompt optimizers reduce that effort by iteratively refining prompts from scored feedback, yet the security of this optimization stage remains underexamined. We present the first systematic analysis of poisoning risks in LLM-based prompt optimization. Using HarmBench, we find systems are substantially more vulnerable to manipulated feedback than to query poisoning alone: feedback-based attacks raise attack success rate (ASR) by up to ΔASR = 0.48. We introduce a simple fake reward attack that requires no access to the reward model and significantly increases vulnerability. We also propose a lightweight highlighting defense that reduces the fake reward ΔASR from 0.23 to 0.07 without degrading utility. These results establish prompt optimization pipelines as a first-class attack surface and motivate stronger safeguards for feedback channels and optimization frameworks.
title Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2510.14381