Saved in:
Bibliographic Details
Main Authors: Yang, Haoming, Ma, Ke, Jia, Xiaojun, Sun, Yingfei, Xu, Qianqian, Huang, Qingming
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.02862
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908424895201280
author Yang, Haoming
Ma, Ke
Jia, Xiaojun
Sun, Yingfei
Xu, Qianqian
Huang, Qingming
author_facet Yang, Haoming
Ma, Ke
Jia, Xiaojun
Sun, Yingfei
Xu, Qianqian
Huang, Qingming
contents Despite the remarkable performance of Large Language Models (LLMs), they remain vulnerable to jailbreak attacks, which can compromise their safety mechanisms. Existing studies often rely on brute-force optimization or manual design, failing to uncover potential risks in real-world scenarios. To address this, we propose a novel jailbreak attack framework, ICRT, inspired by heuristics and biases in human cognition. Leveraging the simplicity effect, we employ cognitive decomposition to reduce the complexity of malicious prompts. Simultaneously, relevance bias is utilized to reorganize prompts, enhancing semantic alignment and inducing harmful outputs effectively. Furthermore, we introduce a ranking-based harmfulness evaluation metric that surpasses the traditional binary success-or-failure paradigm by employing ranking aggregation methods such as Elo, HodgeRank, and Rank Centrality to comprehensively quantify the harmfulness of generated content. Experimental results show that our approach consistently bypasses mainstream LLMs' safety mechanisms and generates high-risk content, providing insights into jailbreak attack risks and contributing to stronger defense strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
Yang, Haoming
Ma, Ke
Jia, Xiaojun
Sun, Yingfei
Xu, Qianqian
Huang, Qingming
Computation and Language
Artificial Intelligence
Despite the remarkable performance of Large Language Models (LLMs), they remain vulnerable to jailbreak attacks, which can compromise their safety mechanisms. Existing studies often rely on brute-force optimization or manual design, failing to uncover potential risks in real-world scenarios. To address this, we propose a novel jailbreak attack framework, ICRT, inspired by heuristics and biases in human cognition. Leveraging the simplicity effect, we employ cognitive decomposition to reduce the complexity of malicious prompts. Simultaneously, relevance bias is utilized to reorganize prompts, enhancing semantic alignment and inducing harmful outputs effectively. Furthermore, we introduce a ranking-based harmfulness evaluation metric that surpasses the traditional binary success-or-failure paradigm by employing ranking aggregation methods such as Elo, HodgeRank, and Rank Centrality to comprehensively quantify the harmfulness of generated content. Experimental results show that our approach consistently bypasses mainstream LLMs' safety mechanisms and generates high-risk content, providing insights into jailbreak attack risks and contributing to stronger defense strategies.
title Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.02862