POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinyu, Huang, Tianjin, Mu, Ronghui, Huang, Xiaowei, Jin, Gaojie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Turning Black Box into White Box: Dataset Distillation Leaks
by: Chen, Huajie, et al.
Published: (2026)
by: Chen, Huajie, et al.
Published: (2026)
Prompt-Induced Over-Generation as Denial-of-Service: A Black-Box Attack-Side Benchmark
by: Manu, et al.
Published: (2025)
by: Manu, et al.
Published: (2025)
Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models
by: Wang, Shuqiang, et al.
Published: (2026)
by: Wang, Shuqiang, et al.
Published: (2026)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
by: Liu, Shuaitong, et al.
Published: (2025)
by: Liu, Shuaitong, et al.
Published: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
by: Sitawarin, Chawin, et al.
Published: (2024)
by: Sitawarin, Chawin, et al.
Published: (2024)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
by: Bokobza, Roey, et al.
Published: (2024)
by: Bokobza, Roey, et al.
Published: (2024)
SurvAttack: Black-Box Attack On Survival Models through Ontology-Informed EHR Perturbation
by: Kerdabadi, Mohsen Nayebi, et al.
Published: (2024)
by: Kerdabadi, Mohsen Nayebi, et al.
Published: (2024)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
by: Yang, Junxiao, et al.
Published: (2025)
by: Yang, Junxiao, et al.
Published: (2025)
Safeguarding Large Language Models: A Survey
by: Dong, Yi, et al.
Published: (2024)
by: Dong, Yi, et al.
Published: (2024)
Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization
by: Guan, Jiwei, et al.
Published: (2026)
by: Guan, Jiwei, et al.
Published: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
by: Chen, Wenyu, et al.
Published: (2026)
by: Chen, Wenyu, et al.
Published: (2026)
Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
by: Guo, Qiming, et al.
Published: (2025)
by: Guo, Qiming, et al.
Published: (2025)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
by: Basani, Advik Raj, et al.
Published: (2024)
by: Basani, Advik Raj, et al.
Published: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Invisible Backdoor Attack Through Singular Value Decomposition
by: Chen, Wenmin, et al.
Published: (2024)
by: Chen, Wenmin, et al.
Published: (2024)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
Black Box Absorption: LLMs Undermining Innovative Ideas
by: Cao, Wenjun
Published: (2025)
by: Cao, Wenjun
Published: (2025)
Principal Eigenvalue Regularization for Improved Worst-Class Certified Robustness of Smoothed Classifiers
by: Jin, Gaojie, et al.
Published: (2025)
by: Jin, Gaojie, et al.
Published: (2025)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026)
by: Wang, Erchi, et al.
Published: (2026)
Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models
by: Wang, Jeffrey G., et al.
Published: (2024)
by: Wang, Jeffrey G., et al.
Published: (2024)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022)
by: Panda, Ashwinee, et al.
Published: (2022)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
by: Gubri, Martin, et al.
Published: (2024)
by: Gubri, Martin, et al.
Published: (2024)
GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs
by: Das, Sanjay, et al.
Published: (2024)
by: Das, Sanjay, et al.
Published: (2024)
A White-Box Adversarial Attack Against a Digital Twin
by: Patterson, Wilson, et al.
Published: (2022)
by: Patterson, Wilson, et al.
Published: (2022)
Differentially Private Iterative Screening Rules for Linear Regression
by: Khanna, Amol, et al.
Published: (2025)
by: Khanna, Amol, et al.
Published: (2025)
Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval
by: Wang, Wenshuo, et al.
Published: (2025)
by: Wang, Wenshuo, et al.
Published: (2025)
ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
by: Li, Yunzhe, et al.
Published: (2025)
by: Li, Yunzhe, et al.
Published: (2025)
BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models
by: Wang, Xinyuan, et al.
Published: (2024)
by: Wang, Xinyuan, et al.
Published: (2024)
LoRID: Low-Rank Iterative Diffusion for Adversarial Purification
by: Zollicoffer, Geigh, et al.
Published: (2024)
by: Zollicoffer, Geigh, et al.
Published: (2024)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
by: Lai, Zhenglin, et al.
Published: (2025)
by: Lai, Zhenglin, et al.
Published: (2025)
FreeMark: A Non-Invasive White-Box Watermarking for Deep Neural Networks
by: Chen, Yuzhang, et al.
Published: (2024)
by: Chen, Yuzhang, et al.
Published: (2024)
Data-Free Privacy-Preserving for LLMs via Model Inversion and Selective Unlearning
by: Zhou, Xinjie, et al.
Published: (2026)
by: Zhou, Xinjie, et al.
Published: (2026)
Differentially Private Worst-group Risk Minimization
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Privacy-Preserving Federated Learning via Homomorphic Adversarial Networks
by: Dong, Wenhan, et al.
Published: (2024)
by: Dong, Wenhan, et al.
Published: (2024)
Similar Items
-
Turning Black Box into White Box: Dataset Distillation Leaks
by: Chen, Huajie, et al.
Published: (2026) -
Prompt-Induced Over-Generation as Denial-of-Service: A Black-Box Attack-Side Benchmark
by: Manu, et al.
Published: (2025) -
Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models
by: Wang, Shuqiang, et al.
Published: (2026) -
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
by: Liu, Shuaitong, et al.
Published: (2025) -
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)