Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chiyu, Zhou, Lu, Xu, Xiaogang, Wu, Jiafei, Fang, Liming, Liu, Zhe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908750440300544
author Zhang, Chiyu
Zhou, Lu
Xu, Xiaogang
Wu, Jiafei
Fang, Liming
Liu, Zhe
author_facet Zhang, Chiyu
Zhou, Lu
Xu, Xiaogang
Wu, Jiafei
Fang, Liming
Liu, Zhe
contents Existing black-box jailbreak attacks achieve certain success on non-reasoning models but degrade significantly on recent SOTA reasoning models. To improve attack ability, inspired by adversarial aggregation strategies, we integrate multiple jailbreak tricks into a single developer template. Especially, we apply Adversarial Context Alignment to purge semantic inconsistencies and use NTP (a type of harmful prompt) -based few-shot examples to guide malicious outputs, lastly forming DH-CoT attack with a fake chain of thought. In experiments, we further observe that existing red-teaming datasets include samples unsuitable for evaluating attack gains, such as BPs, NHPs, and NTPs. Such data hinders accurate evaluation of true attack effect lifts. To address this, we introduce MDH, a Malicious content Detection framework integrating LLM-based annotation with Human assistance, with which we clean data and build RTA dataset suite. Experiments show that MDH reliably filters low-quality samples and that DH-CoT effectively jailbreaks models including GPT-5 and Claude-4, notably outperforming SOTA methods like H-CoT and TAP.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
Zhang, Chiyu
Zhou, Lu
Xu, Xiaogang
Wu, Jiafei
Fang, Liming
Liu, Zhe
Computation and Language
Cryptography and Security
Existing black-box jailbreak attacks achieve certain success on non-reasoning models but degrade significantly on recent SOTA reasoning models. To improve attack ability, inspired by adversarial aggregation strategies, we integrate multiple jailbreak tricks into a single developer template. Especially, we apply Adversarial Context Alignment to purge semantic inconsistencies and use NTP (a type of harmful prompt) -based few-shot examples to guide malicious outputs, lastly forming DH-CoT attack with a fake chain of thought. In experiments, we further observe that existing red-teaming datasets include samples unsuitable for evaluating attack gains, such as BPs, NHPs, and NTPs. Such data hinders accurate evaluation of true attack effect lifts. To address this, we introduce MDH, a Malicious content Detection framework integrating LLM-based annotation with Human assistance, with which we clean data and build RTA dataset suite. Experiments show that MDH reliably filters low-quality samples and that DH-CoT effectively jailbreaks models including GPT-5 and Claude-4, notably outperforming SOTA methods like H-CoT and TAP.
title Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
topic Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2508.10390