Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Jingyu, Wang, Maolin, Wang, Nan, Li, Jiatong, Li, Yuchen, Ye, Yuyang, Wang, Wanyu, Jia, Pengyue, Zhang, Kai, Zhao, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911620173660160
author Peng, Jingyu
Wang, Maolin
Wang, Nan
Li, Jiatong
Li, Yuchen
Ye, Yuyang
Wang, Wanyu
Jia, Pengyue
Zhang, Kai
Zhao, Xiangyu
author_facet Peng, Jingyu
Wang, Maolin
Wang, Nan
Li, Jiatong
Li, Yuchen
Ye, Yuyang
Wang, Wanyu
Jia, Pengyue
Zhang, Kai
Zhao, Xiangyu
contents Despite substantial advancements in aligning large language models (LLMs) with human values, current safety mechanisms remain susceptible to jailbreak attacks. We hypothesize that this vulnerability stems from distributional discrepancies between alignment-oriented prompts and malicious prompts. To investigate this, we introduce LogiBreak, a novel and universal black-box jailbreak method that leverages logical expression translation to circumvent LLM safety systems. By converting harmful natural language prompts into formal logical expressions, LogiBreak exploits the distributional gap between alignment data and logic-based inputs, preserving the underlying semantic intent and readability while evading safety constraints. We evaluate LogiBreak on a multilingual jailbreak dataset spanning three languages, demonstrating its effectiveness across various evaluation settings and linguistic contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13527
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
Peng, Jingyu
Wang, Maolin
Wang, Nan
Li, Jiatong
Li, Yuchen
Ye, Yuyang
Wang, Wanyu
Jia, Pengyue
Zhang, Kai
Zhao, Xiangyu
Computation and Language
Artificial Intelligence
Despite substantial advancements in aligning large language models (LLMs) with human values, current safety mechanisms remain susceptible to jailbreak attacks. We hypothesize that this vulnerability stems from distributional discrepancies between alignment-oriented prompts and malicious prompts. To investigate this, we introduce LogiBreak, a novel and universal black-box jailbreak method that leverages logical expression translation to circumvent LLM safety systems. By converting harmful natural language prompts into formal logical expressions, LogiBreak exploits the distributional gap between alignment data and logic-based inputs, preserving the underlying semantic intent and readability while evading safety constraints. We evaluate LogiBreak on a multilingual jailbreak dataset spanning three languages, demonstrating its effectiveness across various evaluation settings and linguistic contexts.
title Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.13527