Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mustafa, Ahmed B, Ye, Zihan, Lu, Yang, Pound, Michael P, Gowda, Shreyank N
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915416915312640
author Mustafa, Ahmed B
Ye, Zihan
Lu, Yang
Pound, Michael P
Gowda, Shreyank N
author_facet Mustafa, Ahmed B
Ye, Zihan
Lu, Yang
Pound, Michael P
Gowda, Shreyank N
contents Despite significant advancements in alignment and content moderation, large language models (LLMs) and text-to-image (T2I) systems remain vulnerable to prompt-based attacks known as jailbreaks. Unlike traditional adversarial examples requiring expert knowledge, many of today's jailbreaks are low-effort, high-impact crafted by everyday users with nothing more than cleverly worded prompts. This paper presents a systems-style investigation into how non-experts reliably circumvent safety mechanisms through techniques such as multi-turn narrative escalation, lexical camouflage, implication chaining, fictional impersonation, and subtle semantic edits. We propose a unified taxonomy of prompt-level jailbreak strategies spanning both text-output and T2I models, grounded in empirical case studies across popular APIs. Our analysis reveals that every stage of the moderation pipeline, from input filtering to output validation, can be bypassed with accessible strategies. We conclude by highlighting the urgent need for context-aware defenses that reflect the ease with which these jailbreaks can be reproduced in real-world settings.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21820
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is
Mustafa, Ahmed B
Ye, Zihan
Lu, Yang
Pound, Michael P
Gowda, Shreyank N
Computer Vision and Pattern Recognition
Despite significant advancements in alignment and content moderation, large language models (LLMs) and text-to-image (T2I) systems remain vulnerable to prompt-based attacks known as jailbreaks. Unlike traditional adversarial examples requiring expert knowledge, many of today's jailbreaks are low-effort, high-impact crafted by everyday users with nothing more than cleverly worded prompts. This paper presents a systems-style investigation into how non-experts reliably circumvent safety mechanisms through techniques such as multi-turn narrative escalation, lexical camouflage, implication chaining, fictional impersonation, and subtle semantic edits. We propose a unified taxonomy of prompt-level jailbreak strategies spanning both text-output and T2I models, grounded in empirical case studies across popular APIs. Our analysis reveals that every stage of the moderation pipeline, from input filtering to output validation, can be bypassed with accessible strategies. We conclude by highlighting the urgent need for context-aware defenses that reflect the ease with which these jailbreaks can be reproduced in real-world settings.
title Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.21820