POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xinyu, Huang, Tianjin, Mu, Ronghui, Huang, Xiaowei, Jin, Gaojie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915465487450112
author Li, Xinyu
Huang, Tianjin
Mu, Ronghui
Huang, Xiaowei
Jin, Gaojie
author_facet Li, Xinyu
Huang, Tianjin
Mu, Ronghui
Huang, Xiaowei
Jin, Gaojie
contents Recent advances in Chain-of-Thought (CoT) prompting have substantially enhanced the reasoning capabilities of large language models (LLMs), enabling sophisticated problem-solving through explicit multi-step reasoning traces. However, these enhanced reasoning processes introduce novel attack surfaces, particularly vulnerabilities to computational inefficiency through unnecessarily verbose reasoning chains that consume excessive resources without corresponding performance gains. Prior overthinking attacks typically require restrictive conditions including access to external knowledge sources for data poisoning, reliance on retrievable poisoned content, and structurally obvious templates that limit practical applicability in real-world scenarios. To address these limitations, we propose POT (Prompt-Only OverThinking), a novel black-box attack framework that employs LLM-based iterative optimization to generate covert and semantically natural adversarial prompts, eliminating dependence on external data access and model retrieval. Extensive experiments across diverse model architectures and datasets demonstrate that POT achieves superior performance compared to other methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19277
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
Li, Xinyu
Huang, Tianjin
Mu, Ronghui
Huang, Xiaowei
Jin, Gaojie
Machine Learning
Artificial Intelligence
Cryptography and Security
Recent advances in Chain-of-Thought (CoT) prompting have substantially enhanced the reasoning capabilities of large language models (LLMs), enabling sophisticated problem-solving through explicit multi-step reasoning traces. However, these enhanced reasoning processes introduce novel attack surfaces, particularly vulnerabilities to computational inefficiency through unnecessarily verbose reasoning chains that consume excessive resources without corresponding performance gains. Prior overthinking attacks typically require restrictive conditions including access to external knowledge sources for data poisoning, reliance on retrievable poisoned content, and structurally obvious templates that limit practical applicability in real-world scenarios. To address these limitations, we propose POT (Prompt-Only OverThinking), a novel black-box attack framework that employs LLM-based iterative optimization to generate covert and semantically natural adversarial prompts, eliminating dependence on external data access and model retrieval. Extensive experiments across diverse model architectures and datasets demonstrate that POT achieves superior performance compared to other methods.
title POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2508.19277