Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Zhen, Zhang, Zongmin, Liang, Deqi, Sun, Han, Liu, Yule, Shen, Yun, Gao, Xiangshan, Yang, Yilong, Liu, Shuai, Yue, Yutao, He, Xinlei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2511.16278
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908666871939072
author Sun, Zhen
Zhang, Zongmin
Liang, Deqi
Sun, Han
Liu, Yule
Shen, Yun
Gao, Xiangshan
Yang, Yilong
Liu, Shuai
Yue, Yutao
He, Xinlei
author_facet Sun, Zhen
Zhang, Zongmin
Liang, Deqi
Sun, Han
Liu, Yule
Shen, Yun
Gao, Xiangshan
Yang, Yilong
Liu, Shuai
Yue, Yutao
He, Xinlei
contents As LLMs become more common, non-expert users can pose risks, prompting extensive research into jailbreak attacks. However, most existing black-box jailbreak attacks rely on hand-crafted heuristics or narrow search spaces, which limit scalability. Compared with prior attacks, we propose Game-Theory Attack (GTA), an scalable black-box jailbreak framework. Concretely, we formalize the attacker's interaction against safety-aligned LLMs as a finite-horizon, early-stoppable sequential stochastic game, and reparameterize the LLM's randomized outputs via quantal response. Building on this, we introduce a behavioral conjecture "template-over-safety flip": by reshaping the LLM's effective objective through game-theoretic scenarios, the originally safety preference may become maximizing scenario payoffs within the template, which weakens safety constraints in specific contexts. We validate this mechanism with classical game such as the disclosure variant of the Prisoner's Dilemma, and we further introduce an Attacker Agent that adaptively escalates pressure to increase the ASR. Experiments across multiple protocols and datasets show that GTA achieves over 95% ASR on LLMs such as Deepseek-R1, while maintaining efficiency. Ablations over components, decoding, multilingual settings, and the Agent's core model confirm effectiveness and generalization. Moreover, scenario scaling studies further establish scalability. GTA also attains high ASR on other game-theoretic scenarios, and one-shot LLM-generated variants that keep the model mechanism fixed while varying background achieve comparable ASR. Paired with a Harmful-Words Detection Agent that performs word-level insertions, GTA maintains high ASR while lowering detection under prompt-guard models. Beyond benchmarks, GTA jailbreaks real-world LLM applications and reports a longitudinal safety monitoring of popular HuggingFace LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
Sun, Zhen
Zhang, Zongmin
Liang, Deqi
Sun, Han
Liu, Yule
Shen, Yun
Gao, Xiangshan
Yang, Yilong
Liu, Shuai
Yue, Yutao
He, Xinlei
Cryptography and Security
Artificial Intelligence
As LLMs become more common, non-expert users can pose risks, prompting extensive research into jailbreak attacks. However, most existing black-box jailbreak attacks rely on hand-crafted heuristics or narrow search spaces, which limit scalability. Compared with prior attacks, we propose Game-Theory Attack (GTA), an scalable black-box jailbreak framework. Concretely, we formalize the attacker's interaction against safety-aligned LLMs as a finite-horizon, early-stoppable sequential stochastic game, and reparameterize the LLM's randomized outputs via quantal response. Building on this, we introduce a behavioral conjecture "template-over-safety flip": by reshaping the LLM's effective objective through game-theoretic scenarios, the originally safety preference may become maximizing scenario payoffs within the template, which weakens safety constraints in specific contexts. We validate this mechanism with classical game such as the disclosure variant of the Prisoner's Dilemma, and we further introduce an Attacker Agent that adaptively escalates pressure to increase the ASR. Experiments across multiple protocols and datasets show that GTA achieves over 95% ASR on LLMs such as Deepseek-R1, while maintaining efficiency. Ablations over components, decoding, multilingual settings, and the Agent's core model confirm effectiveness and generalization. Moreover, scenario scaling studies further establish scalability. GTA also attains high ASR on other game-theoretic scenarios, and one-shot LLM-generated variants that keep the model mechanism fixed while varying background achieve comparable ASR. Paired with a Harmful-Words Detection Agent that performs word-level insertions, GTA maintains high ASR while lowering detection under prompt-guard models. Beyond benchmarks, GTA jailbreaks real-world LLM applications and reports a longitudinal safety monitoring of popular HuggingFace LLMs.
title "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2511.16278